0% found this document useful (0 votes)
17 views217 pages

Astro320 Notes

These lecture notes cover a comprehensive course on astronomy, focusing on fluid dynamics and radiative processes in astrophysics. The course is divided into three parts: Fluid Dynamics, Collisionless Dynamics, and Radiative Processes, with detailed discussions on various topics such as hydrodynamics, gravitational interactions, and the interaction of light with matter. It is assumed that students have a background in vector calculus, and additional appendices provide further information on relevant mathematical concepts.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views217 pages

Astro320 Notes

These lecture notes cover a comprehensive course on astronomy, focusing on fluid dynamics and radiative processes in astrophysics. The course is divided into three parts: Fluid Dynamics, Collisionless Dynamics, and Radiative Processes, with detailed discussions on various topics such as hydrodynamics, gravitational interactions, and the interaction of light with matter. It is assumed that students have a background in vector calculus, and additional appendices provide further information on relevant mathematical concepts.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

These lecture notes are constantly being updated, extended and improved

This is Version S2025:v1.2


Last Modified: Apr 9, 2025

2
Course Outline
Astronomy is the study of everything in our Universe outside of the Earth’s atmo-
sphere, including the Universe as a whole (‘cosmology’). As such, it involves all
of physics: quantum mechanics, nuclear and particle physics, classical mechanics,
special and general relativity, electromagnetism, statistical physics, hydrodynamics,
plasma physics, and even solid state physics. As we will see, though, with the excep-
tion of rocky planets and asteroids, all objects in the Universe can be characterized
as some kind of fluid. In addition, almost all information we receive from these
objects reaches us in the form of radiation (photons)1 . Consequently, in this course
on physical processes in astronomy we will focus almost exclusively on fluid dy-
namics and radiative processes. Note that we focus exclusively on neutral fluids;
the physics of electrically charged fluids, known as plasmas will not be covered.

We start in Part I with standard hydrodynamics, which applies mainly to neutral,


collisional fluids. We start by discussing the continuity, momentum and energy
equations, for both ideal and non-ideal fluids. Next we discuss a variety of different
flows; vorticity, incompressible barotropic flow, viscous flow, accretion flow, and
turbulent flow, before addressing fluid instabilities and shocks. In Part II, we briefly
focus on collisionless dynamics. We start by discussion potential theory and the
Virial theorem, followed by the Collisionless Boltzmann Equation, from which we
derive the Jeans equations. We then study how these Jeans equations differ from
the Euler equations that describe ideal, collisional fluids, and briefly discuss orbit
theory, integrals of motion, and the Jeans theorem. We end with a treatment of
gravitational interactions among collisionless systems. Finally, in Part III we discuss
radiative processes and the interaction of light with matter (scattering & absorption),
including, among others, Compton scattering, recombination, photo- and collisional
ionization, free-free emission, synchrotron emission, radiative transfer and the Saha
equation.

It is assumed that the student is familiar with vector calculus, with curvi-linear
coordinate systems. A brief overview of these topics is provided in Appendices A-E.
The other appendices present detailed background information that isprovided for
the interested student, but which is not considered part of the course material.

1
Other astrophysical messengers include neutrinos, cosmic rays and gravitational waves

3
CONTENTS
Part I: Fluid Dynamics
1: Introduction to Fluids and Plasmas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
2: Dynamical Treatments of Fluids . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
3: Hydrodynamic Equations for Ideal Fluids . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
4: Viscosity, Conductivity & The Stress Tensor . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
5: Hydrodynamic Equations for Non-Ideal Fluids . . . . . . . . . . . . . . . . . . . . . . . . . 34
6: Equations of State . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
7: Vorticity & Circulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
8; Hydrostatics and Steady Flows . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .53
9: Viscous Flow and Accretion Flow . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
10: Turbulence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .77
11: Sound Waves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
12: Shocks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
13: Fluid Instabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99

Part II: Collisionless Dynamics


14: Potential Theory & the Virial Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110
15: The Collisionless Boltzmann Equation & the Jeans equations . . . . . . . . 118
16: Orbit Theory, Integrals of Motion & the Jeans Theorem . . . . . . . . . . . . . . 127
17: Collisions & Encounters of Collisionless Systems . . . . . . . . . . . . . . . . . . . . . 131

Part III: Radiative Processes


18: Radiation Essentials . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142
19: Thermal Equilibrium & Saha Equation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148
20: The Interaction of Light with Matter. I - Scattering . . . . . . . . . . . . . . . . . . 156
21: The Interaction of Light with Matter. II- Absorption . . . . . . . . . . . . . . . . . 166
22: The Interaction of Light with Matter. III - Extinction . . . . . . . . . . . . . . . .169
23: Radiative Transfer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173
24: Continuum Emission Processes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .181

4
APPENDICES

Appendix A: Vector Calculus . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194


Appendix B: Conservative Vector Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199
Appendix C: Integral Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200
Appendix D: Curvi-Linear Coordinate Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201
Appendix E: The Levi-Civita Symbol . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
Appendix F: The Viscous Stress Tensor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212
Appendix G: The Chemical Potential . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215

5
LITERATURE
The material covered and presented in these lecture notes has relied heavily on a
number of excellent textbooks listed below.

• The Physics of Fluids and Plasmas


by A. Choudhuri (ISBN-0-521-55543)

• The Physics of Astrophysics–I. Radiation


by F. Shu (ISBN-0-935702-64-4)

• The Physics of Astrophysics–II. Gas Dynamics


by F. Shu (ISBN-0-935702-65-2)

• Modern Fluid Dynamics for Physics and Astrophysics


by O. Regev, O. Umurhan & P. Yecko (ISBN-978-1-4939-3163-7)

• Principles of Astrophysical Fluid Dynamics


by C. Clarke & [Link] (ISBN-978-0-470-01306-9)

• Introduction to Modern Magnetohydrodynamics


by S. Galtier (ISBN-978-1-316-69247-9)

• Astrophysics: Decoding the Cosmos


by J. Irwin (ISBN-978-0-470-01306-9)

• Theoretical Astrophysics
by M. Bartelmann (ISBN-978-3-527-41004-0)

• Radiative Processes in Astrophysics


by G Rybicki & A Lightman (ISBN-978-0-471-82759-7)

• Galactic Dynamics
by J. Binney & S. Tremaine (ISBN-978-0-691-13027-9)

• Modern Classical Physics


by K. Thorne & R. Blandford (ISBN-978-0-691-15902-7)

6
7
Part I: Fluid Dynamics

Almost everything we encounter in the Universe, from gas planets to stars, and
from the interstellar medium to galaxies, can be categorized as some kind of fluid.
Hence, understanding astrophysical processes requires a solid understanding of fluid
dynamics. The following chapters present fairly detailed description of the dynamics
of fluids with an application to astrophysics.

Fluid dynamics is a rich topic, and one could easily devote an entire course to it.
The following chapters therefore only scratch the surface of this rich topic. Readers
who want to get more indepth information are referred to the following excellent
textbooks
- The Physics of Fluids and Plasmas by A. Choudhuri
- Modern Fluid Dynamics for Physics and Astrophysics by [Link] et al.
- The Physics of Astrophysics II. Gas Dynamics by F. Shu
- Principles of Astrophysical Fluid Dynamics by C. Clarke & B. Carswell
- Modern Classical Physics by [Link] & R. Blandford

8
CHAPTER 1

Introduction to Fluids & Plasmas

What is a fluid?
A fluid is a substance that can flow, has no fixed shape, and offers little resistance
to an external stress

• In a fluid the constituent particles (atoms, ions, molecules, stars) can ‘freely’
move past one another.

• Fluids take on the shape of their container (if not self-gravitating).

• A fluid changes its shape at a steady rate when acted upon by a stress force.

What is a plasma?
A plasma is a fluid in which (some of) the consistituent particles are electrically
charged, such that the interparticle force (Coulomb force) is long-range in nature.

Fluid Demographics:
All fluids are made up of large numbers of constituent particles, which can be
molecules, atoms, ions, dark matter particles or even stars. Different types of fluids
mainly differ in the nature of their interparticle forces. Examples of inter-particle
forces are the Coulomb force (among charged particles in a plasma), vanderWaals
forces (among molecules in a neutral fluid) and gravity (among the stars in a galaxy).
Fluids can be both collisional or collisionless, where we define a collision as an
interaction between constituent particles that causes the trajectory of at least one
of these particles to be deflected ‘noticeably’. Collisions among particles drive the
system towards thermodynamic equilibrium (at least locally) and the velocity
distribution towards a Maxwell-Boltzmann distribution.

In neutral fluids the particles only interact with each other on very small scales.
Typically the inter-particle force is a vanderWaals force, which drops off very rapidly.
Put differently, the typical cross section for interactions is the size of the particles
(i.e., the Bohr radius for atoms), which is very small. Hence, to good approximation

9
Figure 1: Examples of particle trajectories in (a) a collisional, neutral fluid, (b) a
plasma, and (c) a self-gravitating collisionless, neutral fluid. Note how different the
dynamics are.

particles in a neutral fluid move in straight lines in between highly-localized, large-


angle scattering events (‘collisions’). An example of such a particle trajectory is
shown in Fig. 1a. Unless the fluid is extremely dilute, most neutral fluids are
collisional, meaning that the mean free path of the particles is short compared to the
physical scales of interest. In astrophysics, though, there are cases where this is not
necessarily the case. In such cases, the standard equations of fluid dynamics may
not be valid!

In a fully ionized plasma the particles exert Coulomb forces (F~ ∝ r −2 ) on each other.
Because these are long-range forces, the velocity of a charged particle changes more
likely due to a succession of many small deflections rather than due to one large one.
As a consequence, particles trajectories in a highly ionized plasma (see Fig. 1b) are
very different from those in a neutral fluid.

In a weakly ionized plasma most interactions/collisions are among neutrals or be-


tween neutrals and charged particles. These interactions are short range, and a
weaky ionized plasma therefore behaves very much like a neutral fluid.

In astrophysics we often encounter fluids in which the mean, dominant interparticle


force is gravity. We shall refer to such fluids are N-body systems. Examples are
dark matter halos (if dark matter consists of WIMPs or axions) and galaxies (stars
act like neutral particles exerting gravitational forces on each other). Since gravity
is a long-range force, each particle feels the force from all other particles. Consider
the gravitational force F~i at a position ~xi from all particles in a relaxed, equilibrium

10
system. We can then write that

F~i (t) = hF~ ii + δ F~i (t)

Here hF~ ii is the time (or ensemble) averaged force at the instantaneous position of
particle i and δ F~i (t) is the instantaneous deviation due to the discrete nature of the
particles that make up the system. As N → ∞ then δ F~i → 0 and the system is
said to be collisionless; its dynamics are governed by the collective force from all
particles rather than by collisions/interactions with individual particles.

As you learn in Galactic Dynamics, the relaxation time of a gravitational N-body


system, defined as the time scale on which collisions (i.e., the impact of the δ F~ above)
cause the energies of particles to change considerably, is
N
trelax ≃ tcross
8 ln N
where tcross is the crossing time (comparable to the dynamical time) of the system.
Typically N ∼ 1010 (number of stars in a galaxy) or 1050−60 (number of dark matter
particles in a halo), and tcross is roughly between 1 and 10 percent of the Hubble time
(108 to 109 yr). Hence, the relaxation time is many times the age of the Universe, and
these N-body systems are, for all practical purposes, collisionless. As a consequence,
the particle trajectories are (smooth) orbits (see Fig. 1c), and understanding galactic
dynamics requires therefore a solid understanding of orbits. Put differently, ‘orbits
are the building blocks on galaxies’.

Collisional vs. Collisionless Plasmas: If the collisionality of a gravitational


system just depends on N, doesn’t that mean that plasmas are also collisionless?
After all, the interparticle force in a plasma is the Coulomb force, which has the
same long-range 1/r 2 nature as gravity. And the number of particles N of a typical
plasma is huge (≫ 1010 ) while the dynamical time can be large as well (this obviously
depends on the length scales considered, but these tend to be large for astrophysical
plasmas).
However, an important difference between a gravitational system and a plasma is
that the Coulomb force can be either attractive or repulsive, depending on the elec-
trical charges of the particles. On large scales, plasma are neutral. This charge
neutrality is guaranteed by the fact that any charge imbalance would produce
strong electrostatic forces that quickly re-establish neutrality. As you will learn in
any course on plasma physics (not covered in ASTRO 320), the effect of electrical

11
charges is screened beyond the Debye length:
 1/2
kB T
λD = ≃ 4.9 cm n−1/2 T 1/2
8π n e2

Here n is the number density in cm−3 , T is the temperature in degrees Kelvin, and
e is the electrical charge of an electron in e.s.u. Related to the Debye length is the
Plasma parameter
1
g≡ ≃ 8.6 × 10−3 n1/2 T −3/2
n λ3D

A plasma is (to good approximation) collisionless if the number of particles within


the Debye volume, ND = n λ3D = g −1 is sufficiently large. After all, only those
particles exert Coulomb forces on each other; particles that outside of each others
Debye volume do not exert a long-range Coulomb force on each other.

As an example, let’s consider three different astrophysical plasmas: the ISM (inter-
stellar medium), the ICM (intra-cluster medium), and the interior of the Sun. The
warm phase of the ISM has a temperature of T ∼ 104 K and a number density of
n ∼ 1 cm−3 . This implies ND ∼ 1.2 × 108 . Hence, the warm phase of the ISM can be
treated as a collisionless plasma on sufficiently small time-scales (for example when
treating high-frequency plasma waves). The ICM has a much lower average density
of ∼ 10−4 cm−3 and a much higher temperature (∼ 107 K). This implies a much
larger number of particles per Debye volume of ND ∼ 4 × 1014 . Hence, the ICM can
typically be approximated as a collisionless plasma. The interior of stars, though,
has a similar temperature of ∼ 107 K but at much higher density (n ∼ 1023 cm−3 ),
implying ND ∼ 10. Hence, stellar interiors are highly collisional plasmas!

Magnetohydrodynamics: Plasma are excellent conductors, and therefore are


quickly shorted by currents; hence in many cases one may ignore the electrical field,
and focus exclusively on the magnetic field instead. This is called magneto-hydro-
dynamics, or MHD for short. Many astrophysical plasmas have relatively weak mag-
netic fields, and we therefore don’t make big errors if we ignore them. In this case,
when electromagnetic interactions are not important, plasmas behave very much like
neutral fluids. Because of this, most of the material covered in this course will focus
on neutral fluids, despite the fact that more than 99% of all baryonic matter in the
Universe is a plasma.

12
Compressibility: Fluids and plasmas can be either gaseous or liquid. A gas is
compressible and will completely fill the volume available to it. A liquid, on the
other hand, is (to good approximation) incompressible, which means that a liquid
of given mass occupies a given volume.

NOTE: Although a gas is said to be compressible, many gaseous flows (and virtually
all astrophysical flows) are incompressible. When the gas is in a container, you can
easily compress it with a piston, but if I move my hand (sub-sonically) through the
air, the gas adjust itself to the perturbation in an incompressible fashion (it moves out
of the way at the speed of sound). The small compression at my hand propagates
forward at the speed of sound (sound wave) and disperses the gas particles out of
the way. In astrophysics we rarely encounter containers, and subsonic gas flow is
often treated (to good approximation) as being incompressible.

Throughout what follows, we use ‘fluid’ to mean a neutral fluid, and ‘plasma’ to
refer to a fluid in which the particles are electrically charged.

Ideal (Perfect) Fluids and Ideal Gases:


As we discuss in more detail in Chapter 4, the resistance of fluids to shear distortions
is called viscosity, which is a microscopic property of the fluid that depends on the
nature of its constituent particles, and on thermodynamic properties such as tem-
perature. Fluids are also conductive, in that the microscopic collisions between the
constituent particles cause heat conduction through the fluid. In many fluids encoun-
tered in astrophysics, the viscosity and conduction are very small. An ideal fluid,
also called a perfect fluid, is a fluid with zero viscosity and zero condution.

NOTE: An ideal (or perfect) fluid should NOT be confused with an ideal or perfect
gas, which is defined as a gas in which the pressure is solely due to the kinetic motions
of the constituent particles. As we show in Chapter 6, and as you have probably
seen before, this implies that the pressure can be written as P = n kB T , with n the
particle number density, kB the Boltzmann constant, and T the temperature.

Examples of Fluids in Astrophysics:

• Stars: stars are spheres of gas in hydrostatic equilibrium (i.e., gravitational


force is balanced by pressure gradients). Densities and temperatures in a given

13
star cover many orders of magnitude. To good approximation, its equation
of state is that of an ideal gas.

• Giant (gaseous) planets: Similar to stars, gaseous planets are large spheres
of gas, albeit with a rocky core. Contrary to stars, though, the gas is typically
so dense and cold that it can no longer be described with the equation of
state of an ideal gas.

• Planet atmospheres: The atmospheres of planets are stratified, gaseous flu-


ids retained by the planet’s gravity.

• White Dwarfs & Neutron stars: These objects (stellar remnants) can be
described as fluids with a degenerate equation of state.

• Proto-planetary disks: the dense disks of gas and dust surrounding newly
formed stars out of which planetary systems form.

• Inter-Stellar Medium (ISM): The gas in between the stars in a galaxy.


The ISM is typically extremely complicated, and roughly has a three-phase
structure: it consists of a dense, cold (∼ 10K) molecular phase, a warm
(∼ 104 K) phase, and a dilute, hot (∼ 106 K) phase. Stars form out of the dense
molecular phase, while the hot phase is (shock) heated by supernova explosions.
The reason for this three phase medium is associated with the various cooling
mechanisms. At high temperature when all gas is ionized, the main cooling
channel is Bremmstrahlung (acceleration of free electrons by positively charged
ions). At low temperatures (< 104 K), the main cooling channel is molecular
cooling (or cooling through hyperfine transitions in metals).

• Inter-Galactic Medium (IGM): The gas in between galaxies. This gas


is typically very, very dilute (low density). It is continuously ‘exposed’ to
adiabatic cooling due to the expansion of the Universe, but also is heated by
radiation from stars (galaxies) and AGN (active galactic nuclei). The latter,
called ‘reionization’, assures that the typical temperature of the IGM is ∼
104 K.

• Intra-Cluster Medium (ICM): The hot gas in clusters of galaxies. This is


gas that has been shock heated when it fell into the cluster; typically gas passes
through an accretion shock when it falls into a dark matter halo, converting
its infall velocity into thermal motion.

14
• Accretion disks: Accretion disks are gaseous, viscous disks in which the
viscosity (enhanced due to turbulence) causes a net rate of radial infall towards
the center of the disk, while angular momentum is being transported outwards
(accretion)

• Galaxies (stellar component): as already mentioned above, the stellar com-


ponent of galaxies is a collisionless fluid; to very, very good approximation, two
stars in a galaxy will never experience a direct collision with another star.

• Dark matter halos: Another example of a collisionless fluid (at least, it is


often assumed that dark matter is collisionless)...

15
CHAPTER 2

Dynamical Treatments of Fluids

A dynamical theory consists of two characteristic elements:

1. a way to describe the state of the system

2. a (set of) equation(s) to describe how the state variables change with time

Consider the following examples:

Example 1: a classical dynamical system

This system is described by the position vectors (~x) and momentum vectors (~p) of
all the N particles, i.e., by (~x1 , ~x2 , ..., ~xN , ~p1 , ~p2 , ..., ~pN ).
If the particles are trully classical, in that they can’t emit or absorb radiation, then
one can define a Hamiltonian
N
X
H(~xi , ~pi , t) ≡ H(~x1 , ~x2 , ..., ~xN , ~p1 , p~2 , ..., p~N , t) = ~pi · ~x˙ i − L(~xi , ~x˙ i , t)
i=1

where L(~xi , ~x˙ i , t) is the system’s Lagrangian, and ~x˙ i = d~xi /dt.

The equations that describe the time-evolution of these state-variables are the Hamil-
tonian equations of motion:

∂H ∂H
~x˙ i = ; p~˙ i = −
∂~pi ∂~xi

16
Example 2: an electromagnetic field

~ x) and
The state of this system is described by the electrical and magnetic fields, E(~
~ x), respectively, and the equations that describe their evolution with time are the
B(~
Maxwell equations, which contain the terms ∂ E/∂t ~ ~
and ∂ B/∂t.

Example 3: a quantum system

The state of a quantum system is fully described by the (complex) wavefunction


ψ(~x), the time-evolution of which is described by the Schrödinger equation
∂ψ
ih̄ = Ĥψ
∂t
where Ĥ is now the Hamiltonian operator.

Level Description of state Dynamical equations


0: N quantum particles ψ(~x1 , ~x2 , ..., ~xN ) Schrödinger equation
1: N classical particles (~x1 , ~x2 , ..., ~xN , ~v1 , ~v2 , ..., ~vN ) Hamiltonian equations
2: Distribution function f (~x, ~v, t) Boltzmann equation
3: Continuum model ρ(~x), ~u(~x), P (~x), T (~x) Hydrodynamic equations
Different levels of dynamical theories to describe neutral fluids

The different levels of Fluid Dynamics


There are different ‘levels’ of dynamical theories to describe fluids. Since all fluids
are ultimately made up of constituent particles, and since all particles are ultimately
‘quantum’ in nature, the most ‘basic’ level of fluid dynamics describes the state of
a fluid in terms of the N-particle wave function ψ(~x1 , x~2 , ..., ~xN ), which evolves in
time according to the Schrödinger equation. We will call this the level-0 description
of fluid dynamics. Since N is typically extremely large, this level-0 description is
extremely complicated and utterly unfeasible. Fortunately, it is also unneccesary.

According to what is known as Ehrenfest’s theorem, a system of N quantum


particles can be treated as a system of N classical particles if the characteristic
separation between the particles is large compared to the ‘de Broglie’ wavelength

17
h h
λ= ≃√
p mkB T
Here h is the Planck constant, p is the particle’s momentum, m is the particle mass,
kB is the Boltzman constant, and T is the temperature of the fluid. This de Broglie
wavelength indicates the ‘characteristic’ size of the wave-packet that according to
quantum mechanics describes the particle, and is typically very small. Except for
extremely dense fluids such as white dwarfs and neutron stars, or ‘exotic’ types of
dark matter (i.e., ‘fuzzy dark matter’), the de Broglie wavelength is always much
smaller than the mean particle separation, and classical, Newtonian mechanics suf-
fices. As we have seen above, a classical, Newtonian system of N particles can be
described by a Hamiltonian, and the corresponding equations of motions. We refer
to this as the level-1 description of fluid dynamics (see under ‘example 1’ above).
Clearly, when N is very large is it unfeasible to solve the 2N equations of motion for
all the positions and momenta of all particles. We need another approach.

In the level-2 approach, one introduces the distribution function f (~x, p~, t), which
describes the number density of particles in 6-dimensional ‘phase-space’ (~x, p~) (i.e.,
how many particles are there with positions in the 3D volume ~x +d~x and momenta in
the 3D volume ~p + d~p). The equation that describes how f (~x, p~, t) evolves with time
is called the Boltzmann equation for a neutral fluid. If the fluid is collisionless
this reduces to the Collisionless Boltzmann equation (CBE). If the collisionless
fluid is a plasma, the same equation is called the Vlasov equation. Often the CBE
and the Vlasov equation are used without distinction.

At the final level-3, the fluid is modelled as a continuum. This means we ignore
that fluids are made up of constituent particles, and rather describe the fluid with
continuous fields, such as the density and velocity fields ρ(~x) and ~u(~x) which assign
to each point in space a scalar quantity ρ and a vector quantity ~u, respectively. For
an ideal neutral fluid, the state in this level-3 approach is fully described by four
fields: the density ρ(~x), the velocity field ~u(~x), the pressure P (~x), and the internal,
specific energy ε(~x) (or, equivalently, the temperature T (~x)). In the MHD treat-
ment of plasmas one also needs to specify the magnetic field B(~ ~ x). The equations
that describe the time-evolution of ρ(~x), ~u(~x), and ε(~x) are called the continuity
equation, the Navier-Stokes equations, and the energy equation, respectively.
Collectively, we shall refer to these as the hydrodynamic equations or fluid equa-
tions. In MHD you have to slightly modify the Navier-Stokes equations, and add
an additional induction equation describing the time-evolution of the magnetic

18
field. For an ideal (or perfect) fluid (i.e., no viscosity and/or conductivity), the
Navier-Stokes equations reduce to what are known as the Euler equations. For a
collisionless gravitational system, the equivalent of the Euler equations are called the
Jeans equations.

Throughout this course, we mainly focus on the level-3 treatment, to which we refer
hereafter as the macroscopic approach. However, for completeness we will derive
these continuum equations starting from a completely general, microscopic level-1
treatment. Along the way we will see how subtle differences in the inter-particle forces
gives rise to a rich variety in dynamics (fluid vs. plasma, collisional vs. collisionless).

Fluid Dynamics: The Macroscopic Continuum Approach:


In the macroscopic approach, the fluid is treated as a continuum. It is often useful
to think of this continuum as ‘made up’ of fluid elements (FE). These are small
fluid volumes that nevertheless contain many particles, that are significantly larger
than the mean-free path of the particles, and for which one can define local hydro-
dynamical variables such as density, pressure and temperature. The requirements
are:

1. the FE needs to be much smaller than the characteristic scale in the problem,
which is the scale over which the hydrodynamical quantities Q change by an
order of magnitude, i.e.
Q
lFE ≪ lscale ∼
∇Q
2. the FE needs to be sufficiently large that fluctuations due to the finite number
of particles (‘discreteness noise’) can be neglected, i.e.,
3
n lFE ≫1

where n is the number density of particles.

3. the FE needs to be sufficiently large that it ‘knows’ about the local conditions
through collisions among the constituent particles, i.e.,

lFE ≫ λ

where λ is the mean-free path of the fluid particles.

19
The ratio of the mean-free path, λ, to the characteristic scale, lscale is known as the
Knudsen number: Kn = λ/lscale . Fluids typically have Kn ≪ 1; if not, then one
is not justified in using the continuum approach (level-3) to fluid dynamics, and one
is forced to resort to a more statistical approach (level-2).

Note that fluid elements can NOT be defined for a collisionless fluid (which has
an infinite mean-free path). This is one of the reasons why one cannot use the
macroscopic approach to derive the equations that govern a collisionless fluid.

Fluid Dynamics: closure:


In general, a fluid element is characterized by the following six hydro-dynamical
variables:
mass density ρ [g/cm3 ]
fluid velocity ~u [cm/s] (3 components)
3
pressure P [erg/cm ]
specific internal energy ε [erg/g]
Note that ~u is the velocity of the fluid element, not to be confused with the velocity ~v
of individual fluid particles, used in the Boltzmann distribution function. Rather, ~u
is (roughly) a vector sum of all particles velocities ~v that make up the fluid element.

In the case of an ideal (or perfect) fluid (i.e., with zero viscosity and conductivity),
the Navier-Stokes equations (which are the hydrodynamical momentum equations)
reduce to what are called the Euler equations. In that case, the evolution of fluid
elements is describe by the following set of hydrodynamical equations:

1 continuum equation relating ρ and ~u


3 momentum equations relating ρ, ~u and P
1 energy equation relating ρ, ~u, P and ε
Thus we have a total of 5 equations for 6 unknowns. One can solve the set (‘close
it’) by using a constitutive relation. In almost all cases, this is the equation of
state (EoS) P = P (ρ, ε).

• Sometimes the EoS is expressed as P = P (ρ, T ). In that case another constitution


relation is needed, typically ε = ε(ρ, T ).

20
• If the EoS is barotropic, i.e., if P = P (ρ), then the energy equation is not needed
to close the set of equations. There are two barotropic EoS that are encountered
frequently in astrophysics: the isothermal EoS, which describes a fluid for which
cooling and heating always balance each other to maintain a constant temperature,
and the adiabatic EoS, in which there is no net heating or cooling (other than
adiabatic heating or cooling due to the compression or expansion of volume, i.e., the
P dV work). We will discuss these cases in more detail later in the course.

• No EoS exists for a collisionless fluid. Consequently, for a collisionless fluid one
can never close the set of fluid equations, unless one makes a number of simplifying
assumptions (i.e., one postulates various symmetries)

• If the fluid is not ideal, then the momentum equations include terms that contain
the (kinetic) viscosity, ν, and the energy equation includes a term that contains
the conductivity, K. Both ν and K depend on the mean-free path of the constituent
particles and therefore depend on the temperature and collisional cross-section of
the particles. Closure of the set of hydrodynamic equations then demands additional
constitutive equations ν(T ) and K(T ). Often, though, ν and K are simply assumed
to be constant (the T -dependence is ignored).

• In the case the fluid is exposed to an external force (i.e., a gravitational or


electrical field), the momentum and energy equations contain an extra force term.

• If the fluid is self-gravitating (which is the case, for example, for stars and
galaxies) there is an additional unknown, the gravitational potential Φ. However,
there is also an additional equation, the Poisson equation relating Φ to ρ, so that
the set of equations remains closed.

• In the case of a plasma, the charged particles give rise to electric and magnetic
fields. Each fluid element now carries 6 additional scalars (Ex , Ey , Ez , Bx , By , Bz ),
and the set of equations has to be complemented with the Maxwell equations that
describe the time evolution of E~ and B.
~

21
Fluid Dynamics: Eulerian vs. Lagrangian Formalism:
One distinguishes two different formalisms for treating fluid dynamics:

• Eulerian Formalism: in this formalism one solves the fluid equations ‘at
fixed positions’: the evolution of a quantity Q is described by the local (or
partial, or Eulerian) derivative ∂Q/∂t. An Eulerian hydrodynamics code is a
‘grid-based code’, which solves the hydro equations on a fixed grid, or using
an adaptive grid, which refines resolution where needed. The latter is called
Adaptive Mesh Refinement (AMR).

• Lagrangian Formalism: in this formalism one solves the fluid equations


‘comoving with the fluid’, i.e., either at a fixed particle (collisionless fluid)
or at a fixed fluid element (collisional fluid). The evolution of a quantity Q
is described by the substantial (or Lagrangian) derivative dQ/dt (sometimes
written as DQ/Dt). A Lagrangian hydrodynamics code is a ‘particle-based
code’, which solves the hydro equations per simulation particle. Since it needs
to smooth over neighboring particles in order to compute quantities such as
the fluid density, it is called Smoothed Particle Hydrodynamics (SPH).

To derive an expression for the substantial derivative dQ/dt, realize that Q =


Q(t, x, y, z). When the fluid element moves, the scalar quantity Q experiences a
change
∂Q ∂Q ∂Q ∂Q
dQ = dt + dx + dy + dz
∂t ∂x ∂y ∂z
Dividing by dt yields
dQ ∂Q ∂Q ∂Q ∂Q
= + ux + uy + uz
dt ∂t ∂x ∂y ∂z
where we have used that dx/dt = ux , which is the x-component of the fluid velocity
~u, etc. Hence we have that

dQ ∂Q
= + ~u · ∇Q
dt ∂t
~ x, t), it is straightforward
Using a similar derivation, but now for a vector quantity A(~
to show that

22
~
dA ~
∂A
= ~
+ (~u · ∇) A
dt ∂t
which, in index-notation, is written as
dAi ∂Ai ∂Ai
= + uj
dt ∂t ∂xj

Another way to derive the above relation between the Eulerian and Lagrangian
derivatives, is to think of dQ/dt as
 
dQ Q(~x + δ~x, t + δt) − Q(~x, t)
= lim
dt δt→0 δt
Using that
 
~x(t + δt) − ~x(t) δ~x
~u = lim =
δt→0 δt δt
and
 
Q(~x + δ~x, t) − Q(~x, t)
∇Q = lim
x→0
δ~ δ~x
it is straightforward to show that this results in the same expression for the substan-
tial derivative as above.

Kinematic Concepts: Streamlines, Streaklines and Particle Paths:


In fluid dynamics it is often useful to distinguish the following kinematic constructs:

• Streamlines: curves that are instantaneously tangent to the velocity vector


of the flow. Streamlines show the direction a massless fluid element will travel
in at any point in time.

• Streaklines: the locus of points of all the fluid particles that have passed con-
tinuously through a particular spatial point in the past. Dye steadily injected
into the fluid at a fixed point extends along a streakline.

23
Figure 2: Streaklines showing laminar flow across an airfoil; made by injecting dye
at regular intervals in the flow

• Particle paths: (aka pathlines) are the trajectories that individual fluid ele-
ments follow. The direction the path takes is determined by the streamlines of
the fluid at each moment in time.

Only if the flow is steady, which means that all partial time derivatives (i.e., ∂~u/∂t =
∂ρ/∂t = ∂P/∂t) vanish, will streamlines be identical to streaklines be identical to
particle paths. For a non-steady flow, they will differ from each other.

24
CHAPTER 3

Hydrodynamic Equations for Ideal Fluid

Without any formal derivation (this comes later) we now present the hydrodynamic
equations for an ideal, neutral fluid. Note that these equations adopt the level-3
continuum approach discussed in the previous chapter.

Lagrangian Eulerian

dρ ∂ρ
Continuity Eq: = −ρ ∇ · ~u + ∇ · (ρ~u) = 0
dt ∂t
d~u ∇P ∂~u ∇P
Momentum Eqs: =− − ∇Φ + (~u · ∇) ~u = − − ∇Φ
dt ρ ∂t ρ

dε P L ∂ε P L
Energy Eq: = − ∇ · ~u − + ~u · ∇ε = − ∇ · ~u −
dt ρ ρ ∂t ρ ρ

Hydrodynamic equations for an ideal, neutral fluid in gravitational field

NOTE: students should become familiar with switching between the Eulerian and
Lagrangian equations, and between the vector notation shown above and the
index notation. The latter is often easier to work with. When writing down the
index versions, make sure that each term carries the same index, and make use of the
Einstein summation convention. The only somewhat tricky term is the (~u · ∇) ~u-term
in the Eulerian momentum equations, which in index form is given by uj (∂ui /∂xj ),
where i is the index carried by each term of the equation.

Continuity Equation: this equation expresses mass conservation. This is clear


from the Eulerian form, which shows that changing the density at some fixed point
in space requires a converging, or diverging, mass flux at that location.
If a flow is incompressible, then ∇·~u = 0 everywhere, and we thus have that dρ/dt = 0
(i.e., the density of each fluid element is fixed in time as it moves with the flow). If
a fluid is incompressible, than dρ/dt = 0 and we see that the flow is divergence free

25
(∇ · ~u = 0), which is also called solenoidal.

Momentum Equations: these equations simply state than one can accelerate a
fluid element with either a gradient in the pressure, P , or a gradient in the gravita-
tional potential, Φ. Basically these momentum equations are nothing but Newton’s
F~ = m~a applied to a fluid element. In the above form, valid for an inviscid, ideal
fluid, the momentum equations are called the Euler equations.

Energy Equation: the energy equation states that the only way that the specific,
internal energy, ε, of a fluid element can change, in the absence of conduction, is
by adiabatic compression or expansion, which requires a non-zero divergence of the
velocity field (i.e., ∇ · ~u 6= 0), or by radiation (emission or absorption of photons).
The latter is expressed via the net volumetric cooling rate,
dQ
L=ρ =C−H
dt
Here Q is the thermodynamic heat, and C and H are the net volumetric cooling and
heating rates, respectively.

If the ideal fluid is governed by self-gravity (as opposed to, is placed in an external
gravitational field), then one needs to complement the hydrodynamical equations
with the Poisson equation: ∇2 Φ = 4πGρ. In addition, closure requires an addi-
tional constitutive relations in the form of an equation-of-state P = P (ρ, ε). If
the ideal fluid obeys the ideal gas law, then we have the following two constitutive
relations:
kB T 1 kB T
P = ρ, ε=
µ mp γ − 1 µ mp
(see Chapter 6 for details). Here µ is the mean molecular weight of the fluid in units
of the proton mass, mp , and γ is the adiabatic index, which is often taken to be 5/3
as appropriate for a mono-atomic gas.

Especially for the numerical Eulerian treatment of fluids, it is advantageous to write


the hydro equations in conservative form. Let A(~x, t) be some state variable of
the fluid (either scalar or vector). The evolution equation for A is said to be in
conservative form if
∂A
+ ∇ · F~ (A) = S
∂t
26
Here F~ (A) describes the appropriate flux of A and S describes the various sources
and/or sinks of A. The continuity, momentum and energy equations for an ideal
fluid in conservative form are:

∂ρ
+ ∇ · (ρ~u) = 0
∂t
∂ρ~u
+ ∇ · Π = −ρ∇Φ
∂t
∂E ∂Φ
+ ∇ · [(E + P ) ~u] = ρ −L
∂t ∂t

Here
Π = ρ ~u ⊗ ~u + P
is the momentum flux density tensor (of rank 2), and
 
1 2
E=ρ u +Φ+ε
2

is the energy density.

NOTE: In the expression for the momentum flux density tensor A⊗ ~ B ~ is the tensor
product of A ~ and B~ defined such that (A
~ ⊗ B)
~ ij = ai bj (see Appendix A). Hence,
the index-form of the momentum flux density tensor is simply Πij = ρ ui uj + P δij ,
with δij the Kronecker delta function. Note that this expression is ONLY valid for
an ideal fluid; in the next chapter we shall derive a more general expression for the
momentum flux density tensor.

Note also that whereas there is no source or sink term for the density, gradients in the
gravitational field act as a source of momentum, while its time-variability can cause
an increase or decrease in the energy density of the fluid (if the fluid is collisionless,
we call this violent relaxation). Another source/sink term for the energy density
is radiation (emission or absorption of photons).

27
CHAPTER 4

Viscosity, Conductivity & The Stress Tensor

The hydrodynamic equations presented in the previous chapter are only valid for
an ideal fluid, i.e., a fluid without viscosity and conduction. We now examine the
origin of conduction and viscosity, and link the latter to the stress tensor, which
is an important quantity in all of fluid dynamics.

In an ideal fluid, the particles effectively have a mean-free path of zero, such that they
cannot communicate with their neighboring particles. In reality, though, the mean-
free path, λmfp = (nσ)−1 is finite, and particles ”communicate” with each other
through collisions. These collisions cause an exchange of momentum and energy
among the particles involved, acting as a relaxation mechanism. Note that in a
collisionless system the mean-free path is effectively infinite, and there is no two-
body relaxation, only collective relaxation mechanisms (i.e., violent relaxation or
wave-particle interactions).

• When there are gradients in velocity (”shear”) then the collisions among neigh-
boring fluid elements give rise to a net transport of momentum. The collisions
drive the system towards equilibrium, i.e., towards no shear. Hence, the collisions
act as a resistance to shear, which is called viscosity. See Fig. 3 for an illustration.

• When there are gradients in temperature (or, in other words, in specific inter-
nal energy), then the collisions give rise to a net transport of energy. Again,
the collisions drive the system towards equilibrium, in which the gradients vanish,
and the rate at which the fluid can erase a non-zero ∇T is called the (thermal)
conductivity.

The viscosity, µ, and conductivity, K, are called transport coefficients. Expres-


sions for µ and K in terms of the collision cross section, σ, the fluid’s temperature T ,
and the particle mass m, can be derived in a rigorous manner using what is known
as the Chapman-Enskog expansion. This is a fairly complicated topic, that is
outside of the scope of this course. Intertested reader should consult the classical
monographs ”Statistical Mechanics” by K. Huang, or ”The Mathematical Theory of

28
Figure 3: Illustration of origin of viscosity and shear stress. Three neighboring fluids
elements (1, 2 and 3) have different streaming velocities, ~u. Due to the microscopic
motions and collisions (characterized by a non-zero mean free path), there is a net
transfer of momentum from the faster moving fluid elements to the slower moving
fluid elements. This net transfer of momentum will tend to erase the shear in ~u(~x),
and therefore manifests itself as a shear-resistance, known as viscosity. Due to the
transfer of momentum, the fluid elements deform; in our figure, 1 transfers linear
momentum to the top of 2, while 3 extracts linear momentum from the bottom of 2.
Consequently, fluid element 2 is sheared as depicted in the figure at time t+∆t. From
the perspective of fluid element 2, some internal force (from within its boundaries)
has exerted a shear-stress on its bounding surface.

29
Non-uniform Gases” by S. Chapman and T. Cowling. Using the Chapman-Enskog
expansion one finds the following expressions for µ and K:
 1/2
a m kB T 5
µ= , K = cV µ
σ π 2
Here a is a numerical factor that depends on the details of the interparticle forces, σ
is the collisional cross section, and cV is the specific heat (i.e., per unit mass). Thus,
for a given fluid (given σ and m) we basically have that µ = µ(T ) and K = K(T ).

Note that µ ∝ T 1/2 ; viscosity increases with temperature. This only holds for gases!
For liquids we know from experience that viscosity decreases with increasing tem-
perature (think of honey). Since in astrophysics we are mainly concerned with gas,
µ ∝ T 1/2 will be a good approximation for most of what follows.

Now that we have a rough idea of what viscosity (resistance to shear) and conduc-
tivity (resistance to temperature gradients) are, we have to ask how to incorporate
them into our hydrodynamic equations.

Both transport mechanisms relate to the microscopic velocities of the individual


¯
particles. However, the macroscopic continuum approach of fluid dynamics only
deals with the streaming velocities ~u, which represents the velocities of the fluid
elements. In order to link these different velocities we proceed as follows:

Velocity of fluid particles: We split the velocity, ~v , of a fluid particle in a stream-


ing velocity, ~u, and a ‘random’ velocity, w:
~

~v = ~u + w
~
where h~v i = ~u, hwi
~ = 0 and h.i indicates the average over a fluid element. If we
define vi as the velocity in the i-direction, we have that

hvi vj i = ui uj + hwi wj i
These different velocities allow us to define a number of different velocity tensors:
Stress Tensor: σij ≡ −ρhwi wj i σ = −ρw
~ ⊗w ~
Momentum Flux Density Tensor: Πij ≡ +ρhvi vj i Π = +ρ~v ⊗ ~v
Ram Pressure Tensor: Σij ≡ +ρui uj Σ = +ρ~u ⊗ ~u

30
which are related according to σ = Σ − Π. Note that each of these tensors is man-
ifest symmetric (i.e., σij = σji , etc.), which implies that they have 6 independent
variables.

Note that the stress tensor is related to the microscopic random motions. These
are the ones that give rise to pressure, viscosity and conductivity! The reason that
~ x, n̂) acting on a
σij is called the stress tensor is that it is related to the stress Σ(~
surface with normal vector n̂ located at ~x according to

Σi (n̂) = σij nj

Here Σi (n̂) is the i-component of the stress acting on a surface with normal n̂, whose
j-component is given by nj . Hence, in general the stress will not necessarily be along
the normal to the surface, and it is useful to decompose the stress in a normal
stress, which is the component of the stress along the normal to the surface, and a
shear stress, which is the component along the tangent to the surface.

To see that fluid elements in general are subjected to shear stress, consider the
following: Consider a flow (i.e., a river) in which we inject a small, spherical blob (a
fluid element) of dye. If the only stress to which the blob is subject is normal stress,
the only thing that can happen to the blob is an overall compression or expansion.
However, from experience we know that the blob of dye will shear into an extended,
‘spaghetti’-like feature; hence, the blob is clearly subjected to shear stress, and this
shear stress is obvisouly related to another tensor called the deformation tensor
∂ui
Tij =
∂xj

which describes the (local) shear in the fluid flow.

Since ∂ui /∂xj = 0 in a static fluid (~u(~x) = 0), we see that in a static fluid the stress
tensor can only depend on the normal stress, which we call the pressure.

Pascal’s law for hydrostatistics: In a static fluid, there is no preferred direction,


and hence the (normal) stress has to be isotropic:

static fluid ⇐⇒ σij = −P δij

31
The minus sign is a consequence of the sign convention of the stress.

Sign Convention: The stress Σ(~ ~ x, n̂) acting at location ~x on a surface with normal
n̂, is exerted by the fluid on the side of the surface to which the normal points, on
the fluid from which the normal points. In other words, a positive stress results in
compression. Hence, in the case of pure, normal pressure, we have that Σ = −P .

Viscous Stress Tensor: The expression for the stress tensor in the case of static
fluid motivates us to write in general

σij = −P δij + τij

where we have introduced a new tensor, τij , which is known as the viscous stress
tensor, or the deviatoric stress tensor.

Since the deviatoric stress tensor, τij , is only non-zero in the presence of shear in the
fluid flow, this suggests that
∂uk
τij = Tijkl
∂xl
where Tijkl is a proportionality tensor of rank four. As described in Appendix F
(which is NOT part of the curriculum for this course), most (astrophysical) fluids
are Newtonian, in that they obey a number of conditions. As detailed in that
appendix, for a Newtonian fluid, the relation between the stress tensor and the
deformation tensor is given by
 
∂ui ∂uj 2 ∂uk ∂uk
σij = −P δij + µ + − δij + η δij
∂xj ∂xi 3 ∂xk ∂xk

Here P is the pressure, δij is the Kronecker delta function, µ is the coefficient
of shear viscosity, and η is the coefficient of bulk viscosity (aka the ‘second
viscosity’). We thus see that for a Newtonian fluid, the stress tensor, despite being
a symmetric tensor of rank two (which implies 6 independent variables), only has
three independent components: P , µ and η.

Let’s take a closer look at these three quantities, starting with the pressure P . To be
exact, P is the thermodynamic equilibrium pressure, and is normally computed
thermodynamically from some equation of state, P = P (ρ, T ). It is related to the

32
translational kinetic energy of the particles when the fluid, in equilibrium, has reached
equipartition of energy among all its degrees of freedom, including (in the case of
molecules) rotational and vibrations degrees of freedom.

In addition to the thermodynamic equilibrium pressure, P , we can also define a


mechanical pressure, Pm , which is purely related to the translational motion of
the particles, independent of whether the system has reached full equipartition of
energy. The mechanical pressure is simply the average normal stress and therefore
follows from the stress tensor according to
1 1
Pm = − Tr(σij ) = − (σ11 + σ22 + σ33 )
3 3

Using the above expression for σij , and using that ∂uk /∂xk = ∇ · ~u (Einstein sum-
mation convention), it is easy to see that

Pm = P − η ∇ · ~u

From this expression it is clear that the bulk viscosity, η, is only non-zero if P 6=
Pm . This, in turn, can only happen if the constituent particles of the fluid have
degrees of freedom beyond position and momentum (i.e., when they are molecules
with rotational or vibrational degrees of freedom). Hence, for a fluid of monoatoms
(ideal gas), η = 0. From the fact that P = Pm + η∇ · ~u it is clear that for an
incompressible flow P = Pm and the value of η is irrelevant; bulk viscosity plays
no role in incompressible fluids or flows. The only time when Pm 6= P is when a
fluid consisting of particles with internal degrees of freedom (e.g., molecules) has just
undergone a large volumetric change (i.e., during a shock). In that case there may
be a lag between the time the translational motions reach equilibrium and the time
when the system reaches full equipartition in energy among all degrees of freedom.
In astrophysics, bulk viscosity can generally be ignored, but be aware that it may
be important in shocks. This only leaves the shear viscosity µ, which describes the
ability of the fluid to resist shear stress via momentum transport resulting from
collisions and the non-zero mean free path of the particles.

33
CHAPTER 5

Hydrodynamic Equations for Non-Ideal Fluid

In the hydrodynamic equations for an ideal fluid presented in Chapter 3 we ignored


both viscosity and conductivity. We now examine how our hydrodynamic equations
change when allowing for these two transport mechanisms.

As we have seen in the previous chapter, the effect of viscosity is captured by the
stress tensor, which is given by
 
∂ui ∂uj 2 ∂uk ∂uk
σij = −P δij + τij = −P δij + µ + − δij + η δij
∂xj ∂xi 3 ∂xk ∂xk
Note that in the limit µ → 0 and η → 0, valid for an ideal fluid, σij = −P δij . This
suggests that we can incorporate viscosity in the hydrodynamic equations by simply
replacing the pressure P with the stress tensor, i.e., P δij → −σij = P δij − τij .

Starting from the Euler equation in Lagrangian index form;


dui ∂P ∂Φ
ρ =− −ρ
dt ∂xi ∂xi
we use that ∂P/∂xi = ∂(P δij )/∂xj , and then make the above substitution to obtain

dui ∂(−P δij ) ∂τij ∂Φ


ρ = + −ρ
dt ∂xj ∂xj ∂xi

These momentum equations are called the Navier-Stokes equations.

It is more common, and more useful, to write out the viscous stress tensor, yielding

    
dui ∂P ∂ ∂ui ∂uj 2 ∂uk ∂ ∂uk ∂Φ
ρ =− + µ + − δij + η −ρ
dt ∂xi ∂xj ∂xj ∂xi 3 ∂xk ∂xi ∂xk ∂xi

34
These are the Navier-Stokes equations (in Lagragian index form) in all their glory,
containing both the shear viscosity term and the bulk viscosity term (the latter
is often ignored).

Note that µ and η are usually functions of density and temperature so that they
have spatial variations. However, it is common to assume that these are suficiently
small so that µ and η can be treated as constants, in which case they can be taken
outside the differentials. In what follows we will make this assumption as well.

The Navier-Stokes equations in Lagrangian vector form are


 
d~u 2 1
ρ = −∇P + µ∇ ~u + η + µ ∇(∇ · ~u) − ρ∇Φ
dt 3
If we ignore the bulk viscosity (η = 0) then this reduces to
 
d~u ∇P 2 1
=− + ν ∇ ~u + ∇(∇ · ~u) − ∇Φ
dt ρ 3

where we have introduced the kinetic viscosity ν ≡ µ/ρ. Note that these equations
reduce to the Euler equations in the limit ν → 0. Also, note that the ∇(∇ · ~u)
term is only significant in the case of flows with variable compression (i.e., viscous
dissipation of accoustic waves or shocks), and can often be ignored. This leaves the
ν∇2~u term as the main addition to the Euler equations. Yet, this simple ‘diffuse’
term (describing viscous momentum diffusion) dramatically changes the charac-
ter of the equation, as it introduces a higher spatial derivative. Hence, additional
boundary conditions are required to solve the equations. When solving problems
with solid boundaries (not common in astrophysics), this condition is typically that
the tangential (or shear) velocity at the boundary vanishes. Although this may sound
ad hoc, it is supported by observation; for example, the blades of a fan collect dust.

Recall that when writing the Navier-Stokes equation in Eulerian form, we have that
d~u/dt → ∂~u/∂t + ~u · ∇~u. It is often useful to rewrite this extra term using the vector
calculus identity
 
~u · ~u
~u · ∇~u = ∇ + (∇ × ~u) × ~u
2

35
Hence, for an irrotational flow (i.e., a flow for which ∇ × ~u = 0), we have that
~u · ∇~u = 12 ∇u2 , where u ≡ |~u|.

Next we move to the energy equation, modifying it so as to account for both


viscosity and conduction. We start from
dε ∂ui
ρ = −P −L
dt ∂xi
(see Chapter 3, and recall that, with the Einstein summation convention, ∂ui /∂xi =
∇ · ~u). As with the momentum equations, we include viscosity by making the trans-
formation −P δij → σij = −P δij + τij , which we do as follows:
∂ui ∂ui ∂ui ∂ui
−P → −P δij → −P δij + τij
∂xi ∂xj ∂xj ∂xj
This allows us to write the energy equation in vector form as

ρ = −P ∇ · ~u + V − L
dt
where
∂ui
V ≡ τik
∂xk
is the rate of viscous dissipation which describes the rate at which the work done
against viscous forces is irreversibly converted into internal energy.

Now that we have added the effect of viscosity, what remains is to add conduction.
We can make progress by realizing that, on the microscopic level, conduction arises
from collisions among the constituent particles, causing a flux in internal energy. The
internal energy density of a fluid element is h 21 ρw 2 i, where w
~ = ~v − ~u is the random
motion of the particle wrt the fluid element (see Chapter 4), and the angle brackets
indicate an ensemble average over the particles that make up the fluid element. Based
on this we see that the conductive flux in the i-direction can be written as
1
Fcond,i = h ρw 2 wi i = hρεwi i
2
From experience we also know that we can write the conductive flux as

F~cond = −K ∇T

36
with K the thermal conductivity.

Next we realize that conduction only causes a net change in the internal energy at
some fixed position if the divergence in the conductive flux (∇· F~cond ) at that position
is non-zero. This suggests that the final form of the energy equation, for a non-ideal
fluid, and in Lagrangian vector form, has to be


ρ = −P ∇ · ~u − ∇ · F~cond + V − L
dt

To summarize, below we list the full set of equations of gravitational, radial hydro-
dynamics (ignoring bulk viscosity) 2 .

Full Set of Equations of Gravitational, Radiative Hydrodynamics


Continuity Eq. = −ρ ∇ · ~u
dt
 
d~u 2 1
Momentum Eqs. ρ = −∇P + µ ∇ ~u + ∇(∇ · ~u) − ρ ∇Φ
dt 3


Energy Eq. ρ = −P ∇ · ~u − ∇ · F~cond − L + V
dt

Poisson Eq. ∇2 Φ = 4πGρ


 1/2
1 mkB T 5
Constitutive Eqs. P = P (ρ, ε) , µ = µ(T ) ∝ , K = K(T ) ≃ µ(T ) cV
σ π 2

∂ui
Diss/Cond/Rad V ≡ τik , Fcond,k = hρεwk i , L≡C −H
∂xk

2
Diss/Cond/Rad stands for Dissipation, Conduction, Radiation

37
CHAPTER 6

Equations of State

Equation of State (EoS): a thermodynamic equation describing the state of matter


under a given set of physical conditions. In what follows we will always write our
EoS in the form P = P (ρ, T ). Other commonly used forms are P = P (ρ, ε) or
P = P (ρ, S). In the latter, S is the entropy.

Closure: The hydrodynamic equations for an ideal fluid (continuity eq, momentum
eqs, and energy eq) are 5 equations with 7 unknowns: (ρ, ~u, P , T (or ε), and Φ.
With the addition of the Poisson equation, which relates ρ and Φ. The seventh
and final equation that ensures closure is the equation of state P = P (ρ, T ). Note
that if the EoS is barotropic, i.e., P = P (ρ), then the continuity, momentum and
Poisson equations for a closed set, and the energy equation is not required.

Ideal Gas: a hypothetical gas that consists of identical point particles (i.e. of zero
volume) that undergo perfectly elastic collisions and for which interparticle forces
can be neglected.
An ideal gas obeys the ideal gas law: P V = N kB T .

Here N is the total number of particles, kB is Boltzmann’s constant, and V is the


volume occupied by the fluid. Using that ρ = N µmp /V , where µ is the mean
molecular weight in units of the proton mass mp , we have that the EoS for an
ideal gas is given by

kB T
P = P (ρ, T ) = ρ
µ mp
NOTE: astrophysical gases are often well described by the ideal gas law. Even for a
fully ionized gas, the interparticle forces (Coulomb force) can typically be neglected
(i.e., the potential energies involved are typically < 10% of the kinetic energies).
Ideal gas law breaks down for dense, and cool gases, such as those present in gaseous
planets.

38
Maxwell-Boltzmann Distribution: the distribution of particle momenta, p~ =
m~v , of an ideal gas follows the Maxwell-Boltzmann distribution.
 3/2  
3 1 p2
P(~p) d p~ = exp − d3 p~
2πmkB T 2mkB T
where p2 = ~p · ~p. This distribution follows from maximizing entropy under the
following assumptions:

1. all magnitudes of velocity are a priori equally likely


2. all directions are equally likely (isotropy)
3. total energy is constrained at a fixed value
4. total number of particles is constrained at a fixed value

Using that E = p2 /2m we thus see that P(~p) ∝ e−E/kB T .

NOTE: if there are temperature gradients in the gas, then the particle momenta only
follow the Maxwell-Boltzmann distribution locally, with T = T (~x) being the local
temperature.

Pressure: pressure arises from (elastic) collisions of particles. A particle hitting a


wall head on with momentum p = mv results in a transfer of momentum to the wall
of 2mv. Using this concept, and assuming isotropy for the particle momenta, it can
be shown that

P = ζ n hEi
where ζ = 2/3 (ζ = 1/3) in the case of a non-relativistic (relativistic) fluid, and
Z ∞
hEi = E P(E) dE
0
is the average, translational energy of the particles. In the case of our ideal (non-
relativistic) fluid,
 2 Z ∞ 2
p p 3
hEi = = P(p) dp = kB T
2m 0 2m 2

39
Hence, we find that the EoS for an ideal gas is indeed given by

2 kB T
P = n hEi = n kB T = ρ
3 µmp

Specific Internal Energy: the internal energy per unit mass for an ideal gas is

hEi 3 kB T
ε= =
µmp 2 µmp
Actually, the above derivation is only valid for a true ‘ideal gas’, in which the particles
are point particles. More generally,

1 kB T
ε=
γ − 1 µmp

where γ is the adiabatic index, which for an ideal gas is equal to γ = (q +5)/(q +3),
with q the internal degrees of freedom of the fluid particles: q = 0 for point particles
(resulting in γ = 5/3), while diatomic particles have q = 2 (at sufficiently low
temperatures, such that they only have rotational, and no vibrational degrees of
freedom). The fact that q = 2 in that case arises from the fact that a diatomic
molecule only has two relevant rotation axes; the third axis is the symmetry axis of
the molecule, along which the molecule has negligible (zero in case of point particles)
moment of inertia. Consequently, rotation around this symmetry axis carries no
energy.

Photon gas: Having discussed the EoS of an ideal gas, we now focus on a gas of
photons. Photons have energy E = hν and momentum p = E/c = hν/c, with h the
Planck constant.

Black Body: an idealized physical body that absorbs all incident radiation. A black
body (BB) in thermal equilibrium emits electro-magnetic radiation called black
body radiation.
The spectral number density distribution of BB photons is given by

8πν 2 1
nγ (ν, T ) = 3 hν/k T −1
c e B

40
which implies a spectral energy distribution

8πhν 3 1
u(ν, T ) = nγ (ν, T ) hν =
c3 ehν/kB T − 1
and thus an energy density of
Z ∞
4σSB 4
u(T ) = u(ν, T ) dν = T ≡ ar T 4
0 c
where

2π 5 kB4
σSB =
15h3 c2
is the Stefan-Boltzmann constant and ar ≃ 7.6 × 10−15 erg cm−3 K−4 is called the
radiation constant.

Radiation Pressure: when the photons are reflected off a wall, or when they
are absorbed and subsequently re-emitted by that wall, they transfer twice their
momentum in the normal direction to that wall. Since photons are relativistic, we
have that the EoS for a photon gas is given by

1 1 1 aT 4
P = n hEi = nγ hhνi = u(T ) =
3 3 3 3
where we have used that u(T ) = nγ hEi.

Quantum Statistics: according to quantum statistics, a collection of many indis-


tinguishable elementary particles in thermal equilibrium has a momentum distri-
bution given by
   −1
3 g E(p) − µ
f (~p) d p~ = 3 exp ±1 d3 p~
h kB T
where the signature ± takes the positive sign for fermions (which have half-integer
spin), in which case the distribution is called the Fermi-Dirac distribution, and
the negative sign for bosons (particles with zero or integer spin), in which case the
distribution is called the Bose-Einstein distribution. The factor g is the spin
degeneracy factor, which expresses the number of spin states the particles can
have (g = 1 for neutrinos, g = 2 for photons and charged leptons, and g = 6

41
for quarks). Finally, µ is called the chemical potential, and is a form of potential
energy that is related (in a complicated way) to the number density and temperature
of the particles (see Appendix G).

Classical limit: In the limit where the mean interparticle separation is much larger
than the de Broglie wavelength of the particles, so that quantum effects (e.g., Heisen-
berg’s uncertainty principle) can be ignored, the above distribution function of mo-
menta can be accurately approximated by the Maxwell-Boltzmann distribution.

Heisenberg’s Uncertainty Principle: ∆x ∆px > h (where h = 6.63×10−27 g cm2 s−1


is Planck’s constant). One interpretation of this quantum principle is that phase-
space is quantized; no particle can be localized in a phase-space element smaller than
the fundamental element

∆x ∆y ∆z ∆px ∆py ∆pz = h3

Pauli Exclusion Principle: no more than one fermion of a given spin state can
occupy a given phase-space element h3 . Hence, for electrons, which have g = 2, the
maximum phase-space density is 2/h3 .

Degeneracy: When compressing and/or cooling a fermionic gas, at some point


all possible low momentum states are occupied. Any further compression therefore
results in particles occupying high (but the lowest available) momentum states. Since
particle momentum is ultimately responsible for pressure, this degeneracy manifests
itself as an extremely high pressure, known as degeneracy pressure.

Fermi Momentum: Consider a fully degenerate gas of electrons of electron


density ne . It will have fully occupied the part of phase-space with momenta p ≤
pF . Here pF is the maximum momentum of the particles, and is called the Fermi
momentum. The energy corresponding to the Fermi momentum is called the Fermi
energy, EF and is equal to p2F /2m in the case of a non-relativistic gas, and pF c in
the case of a relativistic gas.
Let Vx be the volume occupied in configuration space, and Vp = 34 πp3F the volume
occupied in momentum space. If the total number of particles is N, and the gas is

42
fully degenerate, then
N 3
Vx Vp = h
2
Using that ne = N/Vx , we find that
 1/3
3
pF = ne h

EoS of Non-Relativistic, Degenerate Gas: Using the information above, it is


relatively straightforward (see Problem Sets) to compute the EoS for a fully degen-
erate gas. Using that for a non-relativistic fluid E = p2 /2m and P = 32 n hEi, while
degeneracy implies that
Z Z
1 Ef 1 p F p2 2 2 3 p2F
hEi = E N(E) dE = V x 4πp dp =
N 0 N 0 2m h3 5 2m
we obtain that
 2/3
1 3 h2 5/3
P = ρ
20 π m8/3

EoS of Relativistic, Degenerate Gas: In the case of a relativistic, degenerate gas,


we use the same procedure as above. However, this time we have that P = 31 n hEi
while E = p c, which results in
 1/3
1 3 c h 4/3
P = ρ
8 π m4/3

43
White Dwarfs and the Chandrasekhar limit: White dwarfs are the end-states
of stars with mass low enough that they don’t form a neutron star. When the
pressure support from nuclear fusion in a star comes to a halt, the core will start
to contract until degeneracy pressure kicks in. The star consists of a fully ionized
plasma. Assume for simplicity that the plasma consists purely of hydrogen, so that
the number density of protons is equal to that of electrons: np = ne . Because of
equipartition

p2p p2
= e
2mp 2me
p
Since mp ≫ me we have also that pp ≫ pe (in fact pp /pe = mp /me ≃ 43).
Consequently, when cooling or compressing the core of a star, the electrons will
become degenerate well before the protons do. Hence, white dwarfs are held up
against collapse by the degeneracy pressure from electrons. Since the electrons
are typically non-relativistic, the EoS of the white dwarf is: P ∝ ρ5/3 . If the white
dwarf becomes more and more massive (i.e., because it is accreting mass from a
companion star), the Pauli-exclusion principle causes the Fermi momentum, pF , to
increase to relativistic values. This softens the EoS towards P ∝ ρ4/3 . Such an
equation of state is too soft to stabilize the white dwarf against gravitational collapse;
the white dwarf collapses until it becomes a neutron star, at which stage it is
supported against further collapse by the degeneracy pressure from neutrons. This
happens when the mass of the white dwarf reaches Mlim ≃ 1.44M⊙ , the so-called
Chandrasekhar limit.

Non-Relativistic Relativistic
non-degenerate P ∝ ρT P ∝ T4
degenerate P ∝ ρ5/3 P ∝ ρ4/3
Summary of equations of state for different kind of fluids

44
CHAPTER 7

Vorticity & Circulation

Vorticity: The vorticity of a flow is defined as the curl of the velocity field:

vorticity : ~ = ∇ × ~u
w
It is a microscopic measure of rotation (vector) at a given point in the fluid, which
can be envisioned by placing a paddle wheel into the flow. If it spins about its axis
at a rate Ω, then w = |w|
~ = 2Ω.

Circulation: The circulation around a closed contour C is defined as the line integral
of the velocity along that contour:
I Z
circulation : ΓC = ~u · d~l = w ~
~ · dS
C S

where S is an arbitrary surface bounded by C. The circulation is a macroscopic


measure of rotation (scalar) for a finite area of the fluid.

Irrotational fluid: An irrotational fluid is defined as being curl-free; hence, w


~ =0
and therefore ΓC = 0 for any C.

Vortex line: a line that points in the direction of the vorticity vector. Hence, a
vortex line relates to w,
~ as a streamline relates to ~u (cf. Chapter 2).

Vortex tube: a bundle of vortex lines. The circularity of a curve C is proportional


to the number of vortex lines that thread the enclosed area.

In an inviscid fluid the vortex lines/tubes move with the fluid: a vortex line an-
chored to some fluid element remains anchored to that fluid element.

45
Figure 4: Evolution of a vortex tube. Solid dots correspond to fluid elements. Due
to the shear in the velocity field, the vortex tube is stretched and tilted. However, as
long as the fluid is inviscid and barotropic Kelvin’s circularity theorem assures that
the circularity is conserved with time. In addition, since vorticity is divergence-free
(‘solenoidal’), the circularity along different cross sections of the same vortex-tube is
the same.

Vorticity equation: The Navier-Stokes momentum equations, in the absence of


bulk viscosity, in Eulerian vector form, are given by
 
∂~u ∇P 2 1
+ (~u · ∇) ~u = − − ∇Φ + ν ∇ ~u + ∇(∇ · ~u)
∂t ρ 3
Using the vector identity (~u · ∇) ~u = 12 ∇u2 + (∇ × ~u) × ~u = ∇(u2 /2) − ~u × w
~ allows
us to rewrite this as
 
∂~u ∇P 1 2 2 1
− ~u × w
~ =− − ∇Φ − ∇u + ν ∇ ~u + ∇(∇ · ~u)
∂t ρ 2 3
If we now take the curl on both sides of this equation, and we use that curl(grad S) =
~ = ∇2 (∇ × A),
0 for any scalar field S, and that ∇ × (∇2 A) ~ we obtain the vorticity
equation:

46
 
∂w
~ ∇P
= ∇ × (~u × w)
~ −∇× + ν∇2 w
~
∂t ρ

~ = ∇S × A
To write this in Lagrangian form, we first use that ∇ × (S A) ~ + S (∇ × A)
~
[see Appendix A] to write

1 1 1 ρ∇(1) − 1∇ρ ∇P × ∇ρ
∇ × ( ∇P ) = ∇( ) × ∇P + (∇ × ∇P ) = × ∇P =
ρ ρ ρ ρ2 ρ2

where we have used, once more, that curl(grad S) = 0. Next, using the vector
identities from Appendix A, we write

∇ × (w
~ × ~u) = w(∇
~ · ~u) − (w
~ · ∇)~u − ~u(∇ · w)
~ + (~u · ∇)w
~
The third term vanishes because ∇ · w
~ = ∇ · (∇ × ~u) = 0. Hence, using that ∂ w/∂t
~ +
(~u · ∇)w
~ = dw/dt
~ we finally can write the vorticity equation in Lagrangian
form:

dw
~ ∇ρ × ∇P
~ · ∇)~u − w(∇
= (w ~ · ~u) + 2
+ ν∇2 w
~
dt ρ

This equation describes how the vorticity of a fluid element evolves with time. We
now describe the various terms of the rhs of this equation in turn:

• (w
~ · ∇)~u: This term represents the stretching and tilting of vortex tubes due
to velocity gradients. To see this, we pick w
~ to be pointing in the z-direction.
Then

∂~u ∂ux ∂uy ∂uz


~ · ∇)~u = wz
(w = wz ~ex + wz ~ey + wz ~ez +
∂z ∂z ∂z ∂z
The first two terms on the rhs describe the tilting of the vortex tube, while the
third term describes the stretching.

• w(∇
~ · ~u): This term describes stretching of vortex tubes due to flow com-
pressibility. This term is zero for an incompressible fluid or flow (∇ · ~u = 0).
Note that, again under the assumption that the vorticity is pointing in the
z-direction,

47
 
∂ux ∂uy ∂uz
w(∇
~ · ~u) = wz + + ~ez
∂x ∂y ∂z

• (∇ρ × ∇P )/ρ2 : This is the baroclinic term. It describes the production of


vorticity due to a misalignment between pressure and density gradients. This
term is zero for a barotropic EoS: if P = P (ρ) the pressure and density
gradiens are parallel so that ∇P × ∇ρ = 0. Obviously, this baroclinic term
also vanishes for an incompressible fluid (∇ρ = 0) or for an isobaric fluid (∇P =
0). The baroclinic term is responsible, for example, for creating vorticity in
pyroclastic flows (see Fig. 5).

• ν∇2 w:
~ This term describes the diffusion of vorticity due to viscosity, and
is obviously zero for an inviscid fluid (ν = 0). Typically, viscosity gener-
ates/creates vorticity at a bounding surface: due to the no-slip boundary con-
dition shear arises giving rise to vorticity, which is subsequently diffused into
the fluid by the viscosity. In the interior of a fluid, no new vorticity is generated;
rather, viscosity diffuses and dissipates vorticity.

• ∇ × F~ : There is a fifth term that can create vorticity, which however does not
appear in the vorticity equation above. The reason is that we assumed that the
only external force is gravity, which is a conservative force and can therefore be
written as the gradient of a (gravitational) potential. More generally, though,
there may be non-conservative, external body forces present, which would give
rise to a ∇ × F~ term in the rhs of the vorticity equation. An example of a non-
conservative force creating vorticity is the Coriolis force, which is responsible
for creating hurricanes.

48
Figure 5: The baroclinic creation of vorticity in a pyroclastic flow. High density fluid
flows down a mountain and shoves itself under lower-density material, thus creating
non-zero baroclinicity.

Using the definition of circulation, it can be shown (here without proof) that
Z  
dΓ ∂w~ ~
= + ∇ × (w~ × ~u) · dS
dt S ∂t
Using the vorticity equation, this can be rewritten as

Z  
dΓ ∇ρ × ∇P
= 2
+ ν∇ w~ + ∇ × F~ · dS
~
dt S ρ2

where, for completeness, we have added in the contribution of an external force F~


(which vanishes if F~ is conservative). Using Stokes’ Curl Theorem (see Appendix B)
we can also write this equation in a line-integral form as
I I I
dΓ ∇P ~
=− · dl + ν ∇ ~u · d~l +
2
F~ · d~l
dt ρ

which is the form that is more often used.

49
NOTE: By comparing the equations expressing dw/dt ~ and dΓ/dt it is clear that
the stretching a tilting terms present in the equation describing dw/dt,
~ are absent
in the equation describing dΓ/dt. This implies that stretching and tilting changes
the vorticity, but keeps the circularity invariant. This is basically the first theorem
of Helmholtz described below.

Kelvin’s Circulation Theorem: The number of vortex lines that thread any
element of area that moves with the fluid (i.e., the circulation) remains unchanged
in time for an inviscid, barotropic fluid, in the absence of non-conservative forces.

The proof of Kelvin’s Circulation Theorem is immediately evident from the


above equation, which shows that dΓ/dt = 0 if the fluid is both inviscid (ν = 0),
barotropic (P = P (ρ) ⇒ ∇ρ × ∇P = 0), and there are no non-conservative forces
(F~ = 0).

We end this chapter on vorticity and circulation with the three theorems of Helmholtz,
which hold in the absence of non-conservative forces (i.e., F~ = 0).

Helmholtz Theorem 1: The strength of a vortex tube, which is defined as the


circularity of the circumference of any cross section of the tube, is constant along its
length. This theorem holds for any fluid, and simply derives from the fact that the
vorticity field is divergence-free (we say solenoidal): ∇ · w ~ = ∇ · (∇ × ~u) = 0. To
see this, use Gauss’ divergence theorem to write that
Z Z
∇·w ~ dV = w~ · d2 S = 0
V S
Here V is the volume of a subsection of the vortex tube, and S is its bounding
surface. Since the vorticity is, by definition, perpendicular to S along the sides of
the tube, the only non-vanishing components to the surface integral come from the
areas at the top and bottom of the vortex tube; i.e.
Z Z Z
2~
~ ·d S =
w ~ · (−n̂) dA +
w w~ · n̂ dA = 0
S A1 A2

where A1 and A2 are the areas of the cross sections that bound the volume V of the
vortex tube. Using Stokes’ curl theorem, we have that

50
Z I
~ · n̂ dA =
w ~u · d~l
A C
Hence we have that ΓC1 = ΓC2 where C1 and C2 are the curves bounding A1 and A2 ,
respectively.

Helmholtz Theorem 2: A vortex line cannot end in a fluid. Vortex lines and tubes
must appear as closed loops, extend to infinity, or start/end at solid boundaries.

Helmholtz Theorem 3: A barotropic, inviscid fluid that is initially irrotational


will remain irrotational in the absence of rotational (i.e., non-conservative) external
forces. Hence, such a fluid does not and cannot create vorticity (except across curved
shocks, see Chapter 10).

The proof of Helmholtz’ third theorem is straightforward. According to Kelvin’s


circulation theorem, a barotropic, inviscid fluid has dΓ/dt = 0 everywhere. Hence,
Z  
dΓ ∂w ~ ~=0
= + ∇ × (w ~ × ~u) · d2 S
dt S ∂t
Since this has to hold for any S, we have that ∂ w/∂t
~ = ∇ × (~u × w).
~ Hence, if w
~ =0
initially, the vorticity remains zero for ever.

51
Figure 6: A beluga whale demonstrating Kelvin’s circulation theorem and Helmholtz’
second theorem by producing a closed vortex tube under water, made out of air.

52
CHAPTER 8

Hydrostatics and Steady Flows

Having derived all the relevant equations for hydrodynamics, we now start examining
several specific flows. Since a fully general solution of the Navier-Stokes equation is
(still) lacking (this is one of the seven Millenium Prize Problems, a solution of which
will earn you $1,000,000), we can only make progress if we make several assumptions.

We start with arguably the simplest possible flow, namely ‘no flow’. This is the area
of hydrostatics in which ~u(~x, t) = 0. And since we seek a static solution, we also
must have that all ∂/∂t-terms vanish. Finally, in what follows we shall also ignore
radiative processes (i.e., we set L = 0).

Applying these restrictions to the continuity, momentum and energy equations (see
box at the end of Chapter 5) yields the following two non-trivial equations:

∇P = −ρ ∇Φ

∇ · F~cond = 0

The first equation is the well known equation of hydrostatic equilibrium, stating
that the gravitational force is balanced by pressure gradients, while the second equa-
tion states that in a static fluid the conductive flux needs to be divergence-free.

To further simplify matters, let’s assume (i) spherical symmetry, and (ii) a barotropic
equation of state, i.e., P = P (ρ).

The equation of hydrostatic equilibrium now reduces to

dP G M(r) ρ(r)
=−
dr r2

53
In addition, if the gas is self-gravitating (such as in a star) then we also have that

dM
= 4πρ(r) r 2
dr

For a barotropic EoS this is a closed set of equations, and the density profile can be
solved for (given proper boundary conditions). Of particular interest in astrophysics,
is the case of a polytropic EoS: P ∝ ρΓ , where Γ is the polytropic index. Note
that Γ = 1 and Γ = γ for isothermal and adiabatic equations of state, respectively.
A spherically symmetric, polytropic fluid in HE is called a polytropic sphere.

Lane-Emden equation: Upon substituting the polytropic EoS in the equation


of hydrostatic equilibrium and using the Poisson equation, one obtains a single
differential equation that completely describes the structure of the polytropic sphere,
known as the Lane-Emden equation:
 
1 d 2 dθ
ξ = −θn
ξ 2 dξ dξ

Here n = 1/(Γ − 1) is related to the polytropic index (in fact, confusingly, some texts
refer to n as the polytropic index),
 1/2
4πGρc
ξ= r
Φ0 − Φc
is a dimensionless radius,
 
Φ0 − Φ(r)
θ=
Φ0 − Φc
with Φc and Φ0 the values of the gravitational potential at the center (r = 0) and
at the surface of the star (where ρ = 0), respectively. The density is related to θ
according to ρ = ρc θn with ρc the central density.

Solutions to the Lane-Emden equation are called polytropes of index n. In general,


the Lane-Emden equation has to be solved numerically subject to the boundary
conditions θ = 1 and dθ/dξ = 0 at ξ = 0. Analytical solutions exist, however, for
n = 0, 1, and 5. Examples of polytropes are stars that are supported by degeneracy
pressure. For example, a non-relativistic, degenerate equation of state has P ∝ ρ5/3

54
(see Chapter 6) and is therefore described by a polytrope of index n = 3/2. In the
relativistic case P ∝ ρ4/3 which results in a polytrope of index n = 3.

Another polytrope that is often encountered in astrophysics is the isothermal


sphere, which has P ∝ ρ and thus n = ∞. It has ρ ∝ r −2 at large radii, which im-
plies an infinite total mass. If one truncates the isothermal sphere at some radius and
embeds it in a medium with external pressure (to prevent the sphere from expand-
ing), it is called a Bonnor-Ebert sphere, which is a structure that is frequently
used to describe molecular clouds.

Stellar Structure: stars are gaseous spheres in hydrostatic equilibrium (except


for radial pulsations, which may be considered perturbations away from HE). The
structure of stars is therefore largely governed by the above equation.
However, in general the equation of state is of the form P = P (ρ, T, {Xi }), where
{Xi } is the set of the abundances of all emements i. The temperature structure of a
star and its abundance ratios are governed by nuclear physics (which provides the
source of energy) and the various heat transport mechanisms.

Heat transport in stars: Typically, ignoring abundance gradients, stars have the
equation of state of an ideal gas, P = P (ρ, T ). This implies that the equations of
stellar structure need to be complemented by an equation of the form

dT
= F (r)
dr
Since T is a measure of the internal energy, the rhs of this equation describes the
heat flux, F (r).

The main heat transport mechanisms in a star are:


• conduction
• convection
• radiation
Note that the fourth heat transport mechanism, advection, is not present in the case
of hydrostatic equilibrium, because ~u = 0.

55
Recall from Chapter 4 that the thermal conductivity K ∝ (kB T )1/2 /σ where σ
is the collisional cross section. Using that kB T ∝ v 2 and that the mean-free path of
the particles is λmfp = 1/(nσ), we have that

K ∝ n λmfp v
with v the thermal, microscopic velocity of the particles (recall that ~u = 0). Since
radiative heat transport in a star is basically the conduction of photons, and since
c ≫ ve and the mean-free part of photons is much larger than that of electrons (after
all, the cross section for Thomson scattering, σT , is much smaller than the typical
cross section for Coulomb interactions), we have that in stars radiation is a far more
efficient heat transport mechanism than conduction. An exception are relativistic,
degenerate cores, for which ve ∼ c and photons and electrons have comparable mean-
free paths.

Convection: convection only occurs if the Schwarzschild Stability Criterion is


violated, which happens when the temperature gradient dT /dr becomes too large
(i.e., larger than the temperature gradient that would exist if the star was adiabatic;
see Chapter 15). If that is the case, convection always dominates over radiation as
the most efficient heat transport mechanism. In general, as a rule of thumb, more
massive stars are more radiative and less convective.

Trivia: On average it takes ∼ 200.000 years for a photon created at the core of the
Sun in nuclear burning to make its way to the Sun’s photosphere; from there it only
takes ∼ 8 minutes to travel to the Earth.

Hydrostatic Mass Estimates: Now let us consider the case of an ideal gas, for
which
kB T
P = ρ,
µmp
but this time the gas is not self-gravitating; rather, the gravitational potential may
be considered ‘external’. A good example is the ICM; the hot gas that permeates
clusters. From the EoS we have that

56
dP ∂P dρ ∂P dT P dρ P dT
= + = +
dr ∂ρ dr ∂T dr ρ dr T dr
   
P r dρ r dT P d ln ρ d ln T
= + = +
r ρ dr T dr r d ln r d ln r
Substitution of this equation in the equation for Hydrostatic equilibrium (HE) yields
 
kB T (r) r d ln ρ d ln T
M(r) = − +
µmp G d ln r d ln r

This equation is often used to measure the ‘hydrostatic’ mass of a galaxy cluster;
X-ray measurements can be used to infer ρ(r) and T (r) (after deprojection, which is
analytical in the case of spherical symmetry). Substitution of these two radial depen-
dencies in the above equation then yields an estimate for the cluster’s mass profile,
M(r). Note, though, that this mass estimate is based on three crucial assump-
tions: (i) sphericity, (ii) hydrostatic equilibrium, and (iii) an ideal-gas EoS. Clusters
typically are not spherical, often are turbulent (such that ~u 6= 0, violating the as-
sumption of HE), and can have significant contributions from non-thermal pressure
due to magnetic fields, cosmic rays and/or turbulence. Including these non-thermal
pressure sources the above equation becomes
 
kB T (r) r d ln ρ d ln T Pnt d ln Pnt
M(r) = − + +
µmp G d ln r d ln r Pth d ln r
were Pnt and Pth are the non-thermal and thermal contributions to the total gas
pressure. Unfortunately, it is extremely difficult to measure Pnt reliably, which is
therefore often ignored. This may result in systematic biases of the inferred cluster
mass (typically called the ‘hydrostatic mass’).

Solar Corona: As a final example of a hydrostatic problem in astrophysics, consider


the problem of constructing a static model for the Solar corona.

The Solar corona is a large, spherical region of hot (T ∼ 106 K) plasma extending
well beyond its photosphere. Let’s assume that the heat is somehow (magnetic
reconnection?) produced in the lower layers of the corona, and try to infer the density,
temperature and pressure profiles under the assumption of hydrostatic equilibrium.

57
We have the boundary condition of the temperature at the base, which we assume
to be T0 = 3 × 106 K, at a radius of r = r0 ∼ R⊙ ≃ 6.96 × 1010 cm. The mass of the
corona is negligble, and we therefore have that

dP G M⊙ µmp P
= − 2
dr r kB T
 
d dT
K r2 = 0
dr dr

where we have used the ideal gas EoS to substitute for ρ. Note that the latter of these
equations follows from ∇ · F~ = 0, which is the energy equation in HE. As we have
seen above K ∝ nλmfp T 1/2 . In a plasma one furthermore has that λmfp ∝ n−1 T 2 ,
which implies that K ∝ T 5/2 . Hence, the second equation can be written as
dT
r 2 T 5/2 = constant
dr
which implies
 −2/7
r
T = T0
r0

Note that this equation satisfies our boundary condition, and that T∞ = limr→∞ T (r) =
0. Substituting this expression for T in the HE equation yields

dP G M⊙ µmp dr
=− 2/7 r 12/7
P kB T0 r0

Solving this ODE under the boundary condition that P = P0 at r = r0 yields


" (  )#
−5/7
7 G M⊙ µmp r
P = P0 exp −1
5 kB T0 r0 r0

Note that
 
7 G M⊙ µmp
lim P = P0 exp − 6 0
=
r→∞ 5 kB T0 r0

58
Hence, you need an external pressure to confine the corona. Well, that seems OK,
given that the Sun is embedded in an ISM, whose pressure we can compute taking
characteristic values for the warm phase (T ∼ 104 K and n ∼ 1 cm−3 ). Note that the
other phases (cold and hot) have the same pressure. Plugging in the numbers, we
find that
P∞ ρ0
∼ 10
PISM ρISM

Since ρ0 ≫ ρISM we thus infer that the ISM pressure falls short, by orders of magni-
tude, to be able to confine the corona....

As first inferred by Parker in 1958, the correct implication of this puzzling result is
that a hydrostatic corona is impossible; instead, Parker made the daring suggestion
that there should be a solar wind, which was observationally confirmed a few years
later.

————————————————-

Having addressed hydrostatics (‘no flow’), we now consider the next simplest flow;
steady flow, which is characterised by ~u(~x, t) = ~u(~x). For steady flow ∂~u/∂t = 0,
and fluid elements move along the streamlines (see Chapter 2).

Using the vector identity (~u · ∇) ~u = 21 ∇u2 + (∇ × ~u) × ~u = ∇(u2 /2) − ~u × w,


~ allows
us to write the Navier-Stokes equation for a steady flow of ideal fluid as
 
u2 ∇P
∇ +Φ + − ~u × w
~ =0
2 ρ

This equation is known as Crocco’s theorem. It is simply another form of the


Euler equation for a static flow. In order to write this in a more ‘useful’ form, we
first proceed to demonstrate that ∇P/ρ an be written in terms of the gradients of
the specific enthalpy, h, and the specific entropy, s:

The enthalpy, H, is a measure for the total energy of a thermodynamic system that
includes the internal energy, U, and the amount of energy required to make room
for it by displacing its environment and establishing its volume and pressure:

59
H = U + PV
The differential of the enthalpy can be written as

dH = dU + P dV + V dP
Using the first law of thermodynamics, according to which dU = dQ − P dV , and
the second law of thermodynamics, according to which dQ = T dS, we can rewrite
this as

dH = T dS + V dP
which, in specific form, becomes
dP
dh = T ds +
ρ
(i.e., we have s = S/m). This relation is one of the Gibbs relations frequently
encountered in thermodynamics. NOTE: for completeness, we point out that this
expression ignores changes in the chemical potential (see Appendix J).

The above expression for dh implies that

∇P
= ∇h − T ∇s
ρ

(for a formal proof, see at the end of this chapter). Now recall from the previous
chapter on vorticity that the baroclinic term is given by
 
∇P ∇ρ × ∇P
∇× =
ρ ρ2

Using the above relation, and using that the curl of the gradient of a scalar vanishes,
we can rewrite this baroclinic term as ∇ × (T ∇s). Now using that ∇ × S A ~ =
∇S × A ~ + S(∇ × A)
~ for a scalar S and a vector A ~ (see Appendix A), we can write
∇×(T ∇s) = ∇T ×∇s. This is another form for the baroclinic term. It demonstrates
that another way to create baroclinicity, and thus vorticity, is by having the gradient
in entropy be misaligned with the gradient in temperature.

60
Using the momentum equation for a steady, ideal fluid, and substituting ∇P/ρ →
∇h − T ∇s, we obtain

∇B = T ∇s + ~u × w
~

where we have introduced the Bernoulli function

u2 u2
B≡ +Φ+h= + Φ + ε + P/ρ
2 2

which obviously is a measure of energy. The above equation is another form of


Crocco’s theorem. It relates entropy gradients to vorticity and gradients in the
Bernoulli function.

Let’s investigate what happens to the Bernoulli function for an ideal fluid in a
steady flow. We start by pointing out that in an ideal fluid there is no conduction
and no dissipation. As a consequence, the flow of an ideal fluid conserves entropy;
ds/dt = 0. We say that the flow is isentropic.

Intermezzo: isentropic vs. adiabatic

We consider a flow to be isentropic if it conserves (specific) entropy,


which implies that ds/dt = 0. Note that an ideal fluid is a fluid without
dissipation (viscosity) and conduction (heat flow). Hence, any flow of
ideal fluid is isentropic. A fluid is said to be isentropic if ∇s = 0. A
process is said to be adiabatic if dQ/dt = 0. Note that, according to
the second law of thermodynamics, T dS ≥ dQ. Equality only holds for
a reversible process; in other words, only if a process is adiabatic and
reversible do we call it isentropic. An irreversible, adiabatic process,
therefore, can still create entropy.

Using that
ds ∂s
= + u · ∇s = 0
dt ∂t

61
we see that for a steady flow of ideal fluid we always have that u · ∇s = 0. In words,
there can’t be an entropy gradient in the direction of the flow.

Using the expression for dh derived above, we see that


dh ds 1 dP
=T +
dt dt ρ dt
and since for an ideal fluid the first term on the rhs vanishes we have that
dh 1 dP
=
dt ρ dt
Writing the Lagrangian derivatives in terms of the Eulerian time derivative, and
using that the latter vanish for a steady flow, we thus see that
1
~u · ∇h = ~u · ∇P
ρ
If we now take Crocco’s theorem, and ‘dot’ it with ~u, we obtain that
 2 
u 1
~u · ∇ + Φ + ~u · ∇P − ~u · (~u × w)~ =0
2 ρ
Since ~u × w
~ is perpendicular to ~u the last term vanishes. Using that for an ideal fluid
dh/dt = (1/ρ)(dP/dt), we can rewrite this as
 2 
u
~u · ∇ + Φ + h = ~u · ∇B = 0
2
Thus we see that steady flow of an ideal fluid obeys the following two relations:

~u · ∇s = 0 and ~u · ∇B = 0
Furthermore, using that
dB ∂B
= + ~u · ∇B
dt ∂t
we immediately see that steady flow of ideal fluid obeys

dB
=0
dt

62
Hence, steady flow of an ideal fluid conserves both entropy and the Bernoulli function!
Using the definition of the Bernoulli function we can write this as

dB d~u dΦ ds 1 dP
= ~u · + +T + =0
dt dt dt dt ρ dt
Since ds/dt = 0 for an ideal fluid, we have that if the flow is such that the gravita-
tional potential along the flow doesn’t change significantly (such that dΦ/dt ≃ 0),
we find that

d~u 1 dP
~u · =−
dt ρ dt

This is known as Bernoulli’s theorem, and states that as the speed of a steady flow
increases, the internal pressure of the ideal fluid must decrease (this can be rather
counter-intuitive). Applications of Bernoulli’s theorem discussed in class include the
shower curtain and the pitot tube (a flow measurement device used to measure fluid
flow velocity).

Finally, we mention an important corollary of Crocco’s theorem. We have seen that


a steady flow obeys

∇B − T ∇s = ~u × w
~
Suppose we have an isentropic, irrotational fluid, which means that ∇s = 0 and
~ = 0 everywhere. Then we also have that ∇B = 0; the Bernoulli function is
w
everywhere the same. Now suppose such a flow encounters a shock. As we will
see in Chapter 12, if the shock is adiabatic in that no radiative losses occur, then
conservation of energy implies that the Bernoulli function remains constant across
the shock. However, a shock will increase the entropy of the gas. And if the strength
of the shock varies in the direction perpendicular to the flow (for example, when the
flow hits a curved shock), then behind the shock we have ∇B = 0 but ∇s 6= 0.
This implies that behind the shock we must have created vorticity. This is one of the
most important mechanisms for creating vorticity in astrophysics!

————————————————-

63
Potential flow: The final flow to consider in this chapter is potential flow. Consider
~ ≡ ∇ × ~u = 0 everywhere. This implies that
an irrotational flow, which satisfies w
there is a scalar function, φu (x), such that ~u = ∇φu , which is why φu (x) is called
the velocity potential. The corresponding flow ~u(~x) is called potential flow.

If the fluid is ideal (i.e., ν = K = 0), and barotropic or isentropic, such that the flow
fluid has vanishing baroclinicity, then Kelvin’s circulation theorem assures that
the flow will remain irrotational throughout (no vorticity can be created), provided
that all forces acting on the fluid are conservative.

If the fluid is incompressible, in addition to being irrotational, then we have that


both the curl and the divergence of the velocity field vanish. This implies that

∇ · ~u = ∇2 φu = 0
This is the well known Laplace equation, familiar from electrostatics. Mathemat-
ically, this equation is of the elliptic PDE type which requires well defined boundary
conditions in order for a solution to both exist and be unique. A classical case of
potential flow is the flow around a solid body placed in a large fluid volume. In this
case, an obvious boundary condition is the one stating that the velocity component
perpendicular to the surface of the body at the body (assumed at rest) is zero. This
is called a Neumann boundary condition and is given by
∂φu
= ~n · ∇φu = 0
∂n
with ~n the normal vector. The Laplace equation with this type of boundary condition
constitutes a well-posed problem with a unique solution. An example of potential
flow around a solid body is shown in Fig. 2 in Chapter 2. We will not examine any
specific examples of potential flow, as this means having to solve a Laplace equation,
which is purely a mathematical exersize. We end, though, by pointing out that real
fluids are never perfectly inviscid (ideal fluids don’t exist). And any flow past a
surface involves a boundary layer inside of which viscosity creates vorticity (due to
no-slip boundary condition, which states that the tangential velocity at the surface
of the body must vanish). Hence, potential flow can never fully describe the flow
around a solid body; otherwise one would run into d’Alembert’s paradox which
is that steady potential flow around a body exerts zero force on the body; in other
words, it costs no energy to move a body through the fluid at constant speed. We
know from everyday experience that this is indeed not true. The solution to the

64
paradox is that viscosity created in the boundary layer, and subsequently dissipated,
results in friction.

Although potential flow around an object can thus never be a full description of the
flow, in many cases, the boundary layer is very thin, and away from the boundary
layer the solutions of potential flow still provide an accurate description of the flow.

————————————————-

As promised in the text, we end this chapter by demonstrating that


dP ∇P
dh = T ds + ⇐⇒ ∇h = T ∇s +
ρ ρ

To see this, use that the natural variables of h are the specific entropy, s, and the
pressure P . Hence, h = h(s, P ), and we thus have that
∂h ∂h
dh = ds + dP
∂s ∂P
From a comparison with the previous expression for dh, we see that
∂h ∂h 1
=T, =
∂s ∂P ρ
which allows us to derive

∂h ∂h ∂h
∇h = ~ex + ~ey + ~ez
∂x ∂y ∂z
     
∂h ∂s ∂h ∂P ∂h ∂s ∂h ∂P ∂h ∂s ∂h ∂P
= + ~ex + + ~ey + + ~ez
∂s ∂x ∂P ∂x ∂s ∂y ∂P ∂y ∂s ∂z ∂P ∂z
   
∂h ∂s ∂s ∂s ∂h ∂P ∂P ∂P
= ~ex + ~ey + ~ez + ~ex + ~ey + ~ez
∂s ∂x ∂y ∂z ∂P ∂x ∂y ∂z
1
= T ∇s + ∇P
ρ

which completes our proof.

————————————————-

65
CHAPTER 9

Viscous Flow and Accretion Flow

As we have seen in our discussion on potential flow in the previous chapter, realistic
flow past an object always involves a boundary layer in which viscosity results in
vorticity. Even if the viscosity of the fluid is small, the no-slip boundary condition
typically implies a region where the shear is substantial, and viscocity thus manifests
itself.

In this chapter we examine two examples of viscous flow. We start with a well-
known example from engineering, known as Poiseuille-Hagen flow through a pipe.
Although not really an example of astrophysical flow, it is a good illustration of how
viscosity manifests itself as a consequence of the no-slip boundary condition. The
second example that we consider is viscous flow in a thin accretion disk. This flow,
which was first worked out in detail in a famous paper by Shakura & Sunyaev in
1973, is still used today to describe accretion disks in AGN and around stars.

————————————————-

Pipe Flow: Consider the steady flow of an incompressible viscous fluid through
a circular pipe of radius Rpipe and lenght L. Let ρ be the density of the fluid as
it flows through the pipe, and let ν = µ/ρ be its kinetic viscosity. Since the
flow is incompressible, we have that fluid density will be ρ throughout. If we pick a
Cartesian coordinate system with the z-axis along the symmetry axis of the cylinder,
then the velocity field of our flow is given by

~u = uz (x, y, z) ~ez

In other words, ux = uy = 0.

Starting from the continuity equation


∂ρ
+ ∇ · ρ~u = 0
∂t

66
Figure 7: Poiseuille-Hagen flow of a viscous fluid through a pipe of radius Rpipe and
lenght L.

and using that all partial time-derivatives of a steady flow vanish, we obtain that
∂ρux ∂ρuy ∂ρuz ∂uz
+ + =0 ⇒ =0
∂x ∂y ∂z ∂z
where we have used that ∂ρ/∂z = 0 because of the incompressibility of the flow.
Hence, we can update our velocity field to be ~u = uz (x, y) ~ez .

Next we write down the momentum equations for a steady, incompressible flow,
which are given by
∇P
(~u · ∇)~u = − + ν∇2 ~u − ∇Φ
ρ
In what follows we assume the pipe to be perpendicular to ∇Φ, so that we may
ignore the last term in the above expression. For the x- and y- components of the
momentum equation, one obtains that ∂P/∂x = ∂P/∂y = 0. For the z-component,
we instead have
∂uz 1 ∂P
uz =− + ν∇2 uz
∂z ρ ∂z
Combining this with our result from the continuity equation, we obtain that

1 ∂P
= ν∇2 uz
ρ ∂z

Next we use that ∂P/∂z cannot depend on z; otherwise uz would depend on z, but
according to the continuity equation ∂uz /∂z = 0. This means that the pressure

67
gradient in the z-direction must be constant, which we write as −∆P/L, where ∆P
is the pressure difference between the beginning and end of the pipe, and the minus
sign us used to indicate that the fluid pressure declines as it flows throught the pipe.

Hence, we have that


∆P
∇2 u z = − = constant
ρν L
At this point, it is useful to switch to cylindrical coordinates, (R, θ, z), with the
z-axis as before. Because of the symmetries involved, we have that ∂/∂θ = 0, and
thus the above expression reduces to
 
1 d duz ∆P
R =−
R dR dR ρν L
(see Appendix D). Rewriting this as
1 ∆P
duz = − R dR
2ρν L
and integrating from R to Rpipe using the no-slip boundary condition that uz (Rpipe ) =
0, we finally obtain the flow solution

∆P  2 
uz (R) = Rpipe − R2
4ρ ν L

This solution is called Poiseuille flow or Poiseuille-Hagen flow.

As is evident from the above expression, for a given pressure difference ∆P , the flow
speed u ∝ ν −1 (i.e., a more viscous fluid will flow slower). In addition, for a given
fluid viscosity, applying a larger pressure difference ∆P results in a larger flow speed
(u ∝ ∆P ).

Now let us compute the amount of fluid that flows through the pipe per unit time:
R
Zpipe
π ∆P 4
Ṁ = 2π ρ uz (R) R dR = R
8 ν L pipe
0
Note the strong dependence on the pipe radius; this makes it clear that a clogging of
the pipe has a drastic impact on the mass flow rate (relevant for both arteries and oil-
pipelines). The above expression also gives one a relatively easy method to measure

68
the viscosity of a fluid: take a pipe of known Rpipe and L, apply a pressure difference
∆P across the pipe, and measure the mass flow rate, Ṁ; the above expression allows
one to then compute ν.

The Poiseuille velocity flow field has been experimentally confirmed, but only for
slow flow! When |~u| gets too large (i.e., ∆P is too large), then the flows becomes
irregular in time and space; turbulence develops and |~u| drops due to the enhanced
drag from the turbulence. This will be discussed in more detail in Chapter 12.

————————————————-

Accretion Disks: We now move to a viscous flow that is more relevant for as-
trophysics; accretion flow. Consider a thin accretion disk surrounding an accreting
object of mass M• ≫ Mdisk (such that we may ignore the disk’s self-gravity). Because
of the symmetries involved, we adopt cylindrical coordinates, (R, θ, z), with the
z-axis perpendicular to the disk. We also have that ∂/∂θ is zero, and we set uz = 0
throughout.

We expect uθ to be the main velocity component, with a small uR component repre-


senting the radial accretion flow. We also take the flow to be incompressible.

Let’s start with the continuity equation, which in our case reads
∂ρ 1 ∂
+ (R ρ uR ) = 0
∂t R ∂R
(see Appendix D for how to express the divergence in cylindrical coordinates).

Next up is the Navier-Stokes equations. For now, we only consider the θ-


component, which is given by
∂uθ ∂uθ uθ ∂uθ ∂uθ uR uθ 1 ∂P
+ uR + + uz + =−
∂t ∂R R ∂θ ∂z R ρ ∂θ
 2 2 2

∂ uθ 1 ∂ uθ ∂ uθ 1 ∂uθ 2 ∂uR uθ ∂Φ
+ ν 2
+ 2 2
+ 2
+ + 2 − 2 +
∂R R ∂θ ∂z R ∂R R ∂θ R ∂θ

NOTE: There are several terms in the above expression that may seem ‘surprising’.
The important thing to remember in writing down the equations in curvi-linear

69
coordinates is that operators can also act on unit-direction vectors. For example,
the θ-component of ∇2~u is NOT ∇2 uθ . That is because the operator ∇2 acts on
uR~eR + uθ~eθ + uz~ez , and the directions of ~eR and ~eθ depend on position! The same
holds for the convective operator (~u · ∇) ~u. The full expressions for both cylindrical
and spherical coordinates are written out in Appendix D.

Setting all the terms containing ∂/∂θ and/or uz to zero, the Navier-Stokes equation
simplifies considerably to
   2 
∂uθ ∂uθ uR uθ ∂ uθ ∂ 2 uθ 1 ∂uθ uθ
ρ + uR + =µ + + − 2
∂t ∂R R ∂R2 ∂z 2 R ∂R R
where we have replaced the kinetic viscosity, ν, with µ = νρ.

Integrating over z and writing Z ∞


ρ dz = Σ
−∞
where Σ is the surface density, as well as neglecting variation of ν, uR and uθ with z
(a reasonable approximation), the continuity and Navier-Stokes equation become
∂Σ 1 ∂
+ (R Σ uR ) = 0
 ∂t R ∂R 
∂uθ ∂uθ uR uθ
Σ + uR + = F (µ, R)
∂t ∂R R
where F (µ, R) describes the various viscous terms.

Next we multiply the continuity equation by Ruθ which we can then write as
∂(Σ R uθ ) ∂(Ruθ ) ∂(Σ R uR uθ ) ∂uθ
−Σ + − R Σ uR =0
∂t ∂t ∂R ∂R
Adding this to R times the Navier-Stokes equation, and rearranging terms, yields
∂(Σ R uθ ) ∂(Σ R uR uθ )
+ + Σ uR uθ = G(µ, R)
∂t ∂R
where G(µ, R) = RF (µ). Next we introduce the angular frequency Ω ≡ uθ /R
which allows us to rewrite the above expression as

∂(Σ R2 Ω) 1 ∂ 
+ Σ R3 Ω uR = G(µ, R)
∂t R ∂R

70
Note that Σ R2 Ω = Σ R uθ is the angular momentum per unit surface area. Hence
the above equation describes the evolution of angular momentum in the accretion
disk. It is also clear, therefore, that G(µ, R) must describe the viscous torque on
the disk material, per unit surface area. To derive an expression for it, recall that
Z  2 
∂ uθ 1 ∂uθ uθ
G(µ, R) = R dz µ + − 2
∂R2 R ∂R R

where we have ignored the ∂ 2 uθ /∂z 2 term which is assumed to be small. Using that
µ = νρ and that µ is independent of R and z (this is an assumption that underlies
the Navier-Stokes equation from which we started) we have that
 2 
∂ uθ 1 ∂uθ uθ
G(µ, R) = ν R Σ + − 2
∂R2 R ∂R R
Next we use that uθ = Ω R to write
∂uθ dΩ
= Ω+R
∂R dR
Substituting this in the above expression for G(µ, R) yield
 2
  
2d Ω dΩ 1 ∂ 3 dΩ
G(µ, R) = ν Σ R + 3R = ν ΣR
dR2 dR R ∂R dR

Substituting this expression for the viscous torque in the evolution equation for the
angular momentum per unit surface density, we finally obtain the full set of equations
that govern our thin accretion disk:

 
∂  1 ∂  1 ∂ 3 dΩ
Σ R2 Ω + Σ R 3 Ω uR = ν ΣR
∂t R ∂R R ∂R dR

∂Σ 1 ∂
+ (R Σ uR ) = 0
∂t R ∂R
 1/2
G M•
Ω=
R3

71
These three equations describe the dynamics of a thin, viscous accretion disk. The
third equation indicates that we assume that the fluid is in Keplerian motion around
the accreting object of mass M• . As discussed further below, this is a reasonable
assumption as long as the accretion disk is thin.

Note that the non-zero uR results in a mass inflow rate

Ṁ (R) = −2πΣ R uR

(a positive uR reflects outwards motion).

Now let us consider a steady accretion disk. This implies that ∂/∂t = 0 and
that Ṁ (R) = Ṁ ≡ Ṁ• (the mass flux is constant throughout the disk, otherwise
∂Σ/∂t 6= 0). In particular, the continuity equation implies that

R Σ u R = C1

Using the above expression for the mass inflow rate, we see that

Ṁ•
C1 = −

Similarly, for the Navier-Stokes equation, we have that


dΩ
Σ R 3 Ω uR − ν Σ R 3 = C2
dR
Using the boundary condition that at the radius of the accreting object, R• , the disk
material must be dragged into rigid rotation (a no-slip boundary condition), which
implies that dΩ/dR = 0 at R = R• , we obtain that

Ṁ•
C2 = R•2 Ω• C1 = − (G M• R• )1/2

Substituting this in the above expression, and using that


 1/2
dΩ d G M• 3Ω
= 3
=−
dR dR R 2R

72
we have that
 −1
Ṁ•  2  3 dΩ
νΣ = − R Ω + (G M• R• )1/2 R
2π dR
"  1/2 #
Ṁ• R•
= + 1−
3π R

This shows that the mass inflow rate and kinetic viscosity depend linearly on each
other.

The gravitational energy lost by the inspiraling material is converted into heat. This
is done through viscous dissipation: viscosity robs the disk material of angular
momentum which in turn causes it to spiral in.

We can work out the rate of viscous dissipation using


∂ui
V = τij
∂xj

where we have that the deviatoric stress tensor is


 
∂ui ∂uj 2 ∂uk
τij = µ + − δij
∂xj ∂xi 3 ∂xk

(see Chapter 4). Note that the last term in the above expression vanishes because
the fluid is incompressible, such that
" 2 #
∂ui ∂uj ∂ui
V=µ +
∂xj ∂xi ∂xj

(remember to apply the Einstein summation convention here!).

In our case, using that ∂/∂θ = ∂/∂z = 0 and that uz = 0, the only surviving terms
are
" 2  2 # "  2  2 #
∂uR ∂uθ ∂uR ∂uR ∂uR ∂uθ
V =µ + + =µ 2 +
∂R ∂R ∂R ∂R ∂R ∂R

73
If we make the reasonable assumption that uR ≪ uθ , we can ignore the first term,
such that we finally obtain
 2  2
∂uθ 2 dΩ
V=µ = µR
∂R dR

which expresses the viscous dissipation per unit volume. Note that there is no viscous
dissipation if dΩ/dR = 0, i.e., in the case of solid body rotation. This makes sense
since in that case there is no shear in the disk.

As before, we now proceed by integrating over the z-direction, to obtain


Z  2  2
dE 2 dΩ 2 dΩ
= µR dz = ν Σ R
dt dR dR

Using our expression for νΣ derived above, we can rewrite this as


"  1/2 #  2
dE Ṁ• 2 R• dΩ
= R 1−
dt 3π R dR

Using once more that dΩ/dR = −(3/2)Ω/R, and integrating over the entire disk
yields the accretion luminosity of a thin accretion disk:

Z∞
dE G M• Ṁ•
Lacc ≡ 2π R dR =
dt 2 R•
R•

To put this in perspective, realize that the gravitation energy of mass m at radius
R• is G M• m/ R• . Thus, Lacc is exactly half of the gravitational energy lost due to
the inflow. This obviously begs the question where the other half went...The answer
is simple; it is stored in kinetic energy at the ‘boundary’ radius R• of the accreting
flow.

We end our discussion on accretion disks with a few words of caution. First of
all, our entire derivation is only valid for a thin accretion disk. In a thin disk, the

74
pressure in the disk must be small (otherwise it would puff up). This means that the
∂P/∂R term in the R-component of the Navier-Stokes equation is small compared
to ∂Φ/∂R = GM/R2 . This in turn implies that the gas will indeed be moving on
Keplerian orbits, as we have assumed. If the accretion disk is thick, the situation is
much more complicated, something that will not be covered in this course.

Finally, let us consider the time scale for accretion. As we have seen above, the
energy loss rate per unit surface area is
 2
2 dΩ 9 G M•
ν ΣR = ν
dR 4 R3
We can compare this with the gravitational potential energy of disk material per
unit surface area, which is
G M• Σ
E=
R

This yields an accretion time scale

E 4 R2 R2
tacc ≡ = ∼
dE/dt 9 ν ν

To estimate this time-scale, we first estimate the molecular viscosity. Recall that
ν ∝ λmfpv with v a typical velocity of the fluid particles. In virtually all cases
encountered in astrophysics, we have that the size of the accretion disk, R, is many,
many orders of magnitude larger than λmfp . As a consequence, the corresponding
tacc easily exceeds the Hubble time!

The conclusion is that molecular viscosity is way too small to result in any signif-
icant accretion in objects of astrophysical size. Hence, other source of viscosity are
required, which is a topic of ongoing discussion in the literature. Probably the most
promising candidates are turbulence (in different forms), and the magneto-rotational
instability (MRI). Given the uncertainties involved, it is common practive to simply
write  −1
P 1 dΩ
ν=α
ρ R dR
where α is a ‘free parameter’. A thin accretion disk modelled this way is often called
an alpha-accretion disk. If you wonder what the origin is of the above expression;

75
Figure 8: Image of the central region of NGC 4261 taken with the Hubble Space
Telescope. It reveals a ∼ 100pc scale disk of dust and gas, which happens to be per-
pendicular to a radio jet that emerges from this galaxy. This is an alledged ‘accretion
disk’ supplying fuel to the central black hole in this galaxy.

it simply comes from assuming that the only non-vanishing off-diagonal term of the
stress tensor is taken to be αP (where P is the value along the diagonal of the stress
tensor).

76
CHAPTER 10

Turbulence

Non-linearity: The Navier-Stokes equation is non-linear. This non-linearity arises


from the convective (material) derivative term

1
~u · ∇~u = ∇u2 − ~u × w
~
2
which describes the ”inertial acceleration” and is ultimately responsible for the origin
of the chaotic character of many flows and of turbulence. Because of this non-
linearity, we cannot say whether a solution to the Navier-Stokes equation with nice
and smooth initial conditions will remain nice and smooth for all time (at least not
in 3D).

Laminar flow: occurs when a fluid flows in parallel layers, without lateral mixing
(no cross currents perpendicular to the direction of flow). It is characterized by high
momentum diffusion and low momentum convection.

Turbulent flow: is characterized by chaotic and stochastic property changes. This


includes low momentum diffusion, high momentum convection, and rapid variation
of pressure and velocity in space and time.

The Reynold’s number: In order to gauge the importance of viscosity for a fluid,
it is useful to compare
 2 the ratio of the inertial acceleration (~u · ∇~u) to the viscous
1
acceleration (ν ∇ ~u + 3 ∇(∇ · ~u) ). This ratio is called the Reynold’s number, R,
and can be expressed in terms of the typical velocity scale U ∼ |~u| and length scale
L ∼ 1/∇ of the flow, as

~u · ∇~u U 2 /L UL
R=   ∼ =
ν ∇2~u + 13 ∇(∇ · ~u) νU/L2 ν

If R ≫ 1 then viscosity can be ignored (and one can use the Euler equations to
describe the flow). However, if R ≪ 1 then viscosity is important.

77
Figure 9: Illustration of laminar vs. turbulent flow.

Similarity: Flows with the same Reynold’s number are similar. This is evident
from rewriting the Navier-Stokes equation in terms of the following dimensionless
variables

~u ~x U P Φ ˜ = L∇
ũ = x̃ = t̃ = t p̃ = Φ̃ = ∇
U L L ρ U2 U2
This yields (after multiplying the Navier-Stokes equation with L/U 2 ):
 
∂ ũ ˜ + ∇p̃
˜ +∇˜ Φ̃ = 1 1
˜ ũ + ∇(
2 ˜ ∇
˜ · ũ)
+ ũ · ∇ũ ∇
∂ t̃ R 3

which shows that the form of the solution depends only on R. This principle is
extremely powerful as it allows one to making scale models (i.e., when developing
airplanes, cars etc). NOTE: the above equation is only correct for an incompressible
fluid, i.e., a fluid that obeys ∇ρ = 0. If this is not the case the term P̃ (∇ρ/ρ) needs
to be added at the rhs of the equation, braking its scale-free nature.

78
Figure 10: Illustration of flows at different Reynolds number.

As a specific example, consider fluid flow past a cylinder of diameter L:

• R ≪ 1: ”creeping flow”. In this regime the flow is viscously dominated and


(nearly) symmetric upstream and downstream. The inertial acceleration (~u ·
∇~u) can be neglected, and the flow is (nearly) time-reversible.

• R ∼ 1: Slight asymmetry develops

• 10 ≤ R ≤ 41: Separation occurs, resulting in two counter-rotating votices in


the wake of the cylinder. The flow is still steady and laminar, though.

• 41 ≤ R ≤ 103 : ”von Kármán vortex street”; unsteady laminar flow with


counter-rotating vortices shed periodically from the cylinder (see Fig. 11). Even
at this stage the flow is still ‘predictable’.

• R > 103 : vortices are unstable, resulting in a turbulent wake behind the
cylinder that is ‘unpredictable’.

79
Figure 11: The image shows the von Kármán Vortex street behind a 6.35 mm di-
ameter circular cylinder in water at Reynolds number of 168. The visualization was
done using hydrogen bubble technique. Credit: Sanjay Kumar & George Laughlin,
Department of Engineering, The University of Texas at Brownsville

The following movie shows a R = 250 flow past a cylinder. Initially one can witness
separation, and the creation of two counter-rotating vortices, which then suddenly
become ‘unstable’, resulting in the von Kármán vortex street:
[Link]

80
Figure 12: Typical Reynolds numbers for various biological organisms. Reynolds
numbers are estimated using the length scales indicated, the “rule-of-thumb” in the
text, and material properties of water.

Locomotion at Low-Reynolds number: Low Reynolds number corresponds to


high kinetic visocisity for a given U and L. In this regime of ‘creeping flow’ the
flow past an object is (nearly) time-reversible. Imagine trying to move (swim) in a
highly viscous fluid (take honey as an example). If you try to do so by executing
time-symmetric movements, you will not move. Instead, you need to think of a
symmetry-breaking solution. Nature has found many solutions for this problem.
If we make the simplifying ”rule-of-thumb” assumption that an animal of size L
meters moves roughly at a speed of U = L meters per second (yes, this is very,
very rough, but an ant does move close to 1 mm/s, and a human at roughly 1 m/s),
then we have that R = UL/ν ≃ L2 /ν. Hence, with respect to a fixed substance (say
water, for which ν ∼ 10−2cm2 /s), smaller organisms move at lower Reynolds number
(effectively in a fluid of higher viscosity). Scaling down from a human to bacteria
and single-cell organisms, the motion of the latter in water has R ∼ 10−5 − 10−2
(see Fig. 12). Understanding the locomotion of these organisms is a fascinating
sub-branch of bio-physics.

81
Boundary Layers: Even when R ≫ 1, viscosity always remains important in thin
boundary layers adjacent to any solid surface. This boundary layer must exist in
order to satisfy the no-slip boundary condition. If the Reynolds number exceeds
a critical value, the boundary layer becomes turbulent. Turbulent layers and their
associated turbulent wakes exert a much bigger drag on moving bodies than their
laminar counterparts.

Momentum Diffusion & Reynolds stress: This gives rise to an interesting phe-
nomenon. Consider flow through a pipe. If you increase the viscosity (i.e., decrease
R), then it requires a larger force to achieve a certain flow rate (think of how much
harder it is to push honey through a pipe compared to water). However, this trend
is not monotonic. For sufficiently low viscosity (large R), one finds that the trend
reverses, and that is becomes harder again to push the fluid through the pipe. This
is a consequence of turbulence, which causes momentum diffusion within the flow,
which acts very much like viscosity. However, this momentum diffusion is not due
to the viscous stress tensor, τij , but rather to the Reynolds stress tensor Rij .
To understand the ‘origin’ of the Reynolds stress tensor,consider the following:

For a turbulent flow, ~u(t), it is advantageous to decompose each component of ~u into


a ‘mean’ component, ūi , and a ‘fluctuating’ component, u′i , according to

ui = ūi + u′i
This is knowns as the Reynolds decomposition. The ‘mean’ component can be a
time-average, a spatial average, or an ensemble average, depending on the detailed
characteristics of the flow. Note that this is reminiscent of how we decomposed the
microscopic velocities of the fluid particles in a ‘mean’ velocity (describing the fluid
elements) and a ‘random, microscopic’ velocity (~v = ~u + w).
~

Substituting this into the Navier-Stokes equation, and taking the average of that, we
obtain
∂ ūi ∂ ūi 1 ∂  
+ ūj = σ ij − ρu′i u′j
∂t ∂xj ρ ∂xj
where, for simplicity, we have ignored gravity (the ∇Φ-term). This equation looks
identical to the Navier-Stokes equation (in absence of gravity), except for the −ρu′i u′j
term, which is what we call the Reynolds stress tensor:

82
Rij = −ρu′i u′j
Note that u′i u′j means the same averaging (time, space or ensemble) as above, but
now for the product of u′i and u′j . Note that ū′i = 0, by construction. However,
the expectation value for the product of u′i and u′j is generally not. As is evident
from the equation, the Reynolds stresses (which reflect momentum diffusion due
to turbulence) act in exactly the same way as the viscous stresses. However, they
are only present when the flow is turbulent.

Note also that the Reynolds stress tensor is related to the two-point correlation
tensor

ξij (~r) ≡ u′i (~x, t) u′j (~x + ~r, t)


in the sense that Rij = ξij (0). At large separations, ~r, the fluctuating velocities
will be uncorrelated so that limr→∞ ξij = 0. But on smaller scales the fluctuating
velocities will be correlated, and there will be a ‘characteristic’ scale associated with
these correlations, called the correlation length.

Turbulence: Turbulence is still considered as one of the last ”unsolved problems of


classical physics” [Richard Feynman]. What we technically mean by this is that we
do not yet know how to calculate ξij (~r) (and higher order correlation functions, like
the three-point, four-point, etc) in a particular situation from a fundamental theory.
Salmon (1998) nicely sums up the challenge of defining turbulence:

Every aspect of turbulence is controversial. Even the definition of


fluid turbulence is a subject of disagreement. However, nearly everyone
would agree with some elements of the following description:
• Turbulence requires the presence of vorticity; irrotational flow is
smooth and steady to the extent that the boundary conditions per-
mit.
• Turbulent flow has a complex structure, involving a broad range of
space and time scales.
• Turbulent flow fields exhibit a high degree of apparent randomness
and disorder. However, close inspection often reveals the presence
of embedded cohererent flow structures

83
• Turbulent flows have a high rate of viscous energy dissipation.
• Advected tracers are rapidly mixed by turbulent flows.
However, one further property of turbulence seems to be more fun-
damental than all of these because it largely explains why turbulence
demands a statistical treatment...turbulence is chaotic.

The following is a brief, qualitative description of turbulence:

Turbulence kicks in at sufficiently high Reynolds number (typically R > 103 − 104 ).
Turbulent flow is characterized by irregular and seemingly random motion. Large
vortices (called eddies) are created. These contain a large amount of kinetic energy.
Due to vortex stretching these eddies are stretched thin until they ‘break up’ in
smaller eddies. This results in a cascade in which the turbulent energy is transported
from large scales to small scales. This cascade is largely inviscid, conserving the total
turbulent energy. However, once the length scale of the eddies becomes comparable
to the mean free path of the particles, the energy is dissipated; the kinetic energy
associated with the eddies is transformed into internal energy. The scale at which
this happens is called the Kolmogorov length scale. The length scales between
the scale of turbulence ‘injection’ and the Kolomogorov length scale at which it
is dissipated is called the inertial range. Over this inertial range turbulence is
believed/observed to be scale invariant. The ratio between the injection scale, L,
and the dissipation scale, l, is proportional to the Reynolds number according to
L/l ∝ R3/4 . Hence, two turbulent flows that look similar on large scales (comparable
L), will dissipate their energies on different scales, l, if their Reynolds numbers are
different.

Molecular clouds: an example of turbulence in astrophysics are molecular clouds.


These are gas clouds of masses 105 − 106 M⊙ , densities nH ∼ 100 − 500 cm−3 , and
temperatures T ∼ 10K. They consist mainly of molecular hydrogen and are the
main sites of star formation. Observations show that their velocity linewidths are ∼
6−10km/s, which is much higher than their sound speed (cs ∼ 0.2km/s). Hence, they
are supported against (gravitational) collapse by supersonic turbulence. On small
scales, however, the turbulent motions compress the gas to high enough densities
that stars can form. A numerical simulation of a molecular cloud with supersonic
turbulence is available here:
[Link]

84
CHAPTER 11

Sound Waves

If a (compressible) fluid in equilibrium is perturbed, and the perturbation is suffi-


ciently small, the perturbation will propagate through the fluid as a sound wave
(aka acoustic wave), which is a mechanical, longitudinal wave (i.e, a displacement in
the same direction as that of propagation).

If the perturbation is small, we may assume that the velocity gradients are so small
that viscous effects are negligble (i.e., we can set ν = 0). In addition, we assume that
the time scale for conductive heat transport is large, so that energy exchange due to
conduction can also safely be ignored. In the absence of these dissipative processes,
the wave-induced changes in gas properties are adiabatic.

Before proceeding, let us examine the Reynold’s number of a (propagating) sound


wave. Using that R = U L/ν, and setting U = cs (the typical velocity involved is the
sound speed, to be defined below), L = λ (the characteristic scale of the flow is the
wavelength of the acoustic wave), we have that R = λ cs /ν. Using the expressions
for the viscosity µ = νρ from the constitutive relations in Chapter 5, we see that
ν ∝ λmfp cs . Hence, we have that
UL λ
R≡ ∝
ν λmfp
Thus, as long as the wave-length of the acoustic wave is much larger than the mean-
free path of the fluid particles, we have that the Reynolds number is large, and thus
that viscosity and conduction can be ignored.

Let (ρ0 , P0 , ~u0) be a uniform, equilibrium solution of the Euler fluid equations
(i.e., ignore viscosity). Also, in what follows we will ignore gravity (i.e., ∇Φ = 0).

Uniformity implies that ∇ρ0 = ∇P0 = ∇~u0 = 0. In addition, since the only al-
lowed motion is uniform motion of the entire system, we can always use a Galilean
coordinate transformation so that ~u0 = 0, which is what we adopt in what follows.

85
Substitution into the continuity and momentum equations, one obtains that ∂ρ0 /∂t =
∂~u0 /∂t = 0, indicative of an equilibrium solution as claimed.

Perturbation Analysis: Consider a small perturbation away from the above equi-
librium solution:
ρ0 → ρ0 + ρ1
P0 → P0 + P1
~u0 → ~u0 + ~u1 = ~u1
where |ρ1 /ρ0 | ≪ 1, |P1 /P0 | ≪ 1 and ~u1 is small (compared to the sound speed, to
be derived below).

Substitution in the continuity and momentum equations yields


∂(ρ0 + ρ1 )
+ ∇(ρ0 + ρ1 )~u1 = 0
∂t
∂~u1 ∇(P0 + P1 )
+ ~u1 · ∇~u1 = −
∂t (ρ0 + ρ1 )
which, using that ∇ρ0 = ∇P0 = ∇~u0 = 0 reduces to
∂ρ1
+ ρ0 ∇~u1 + ∇(ρ1~u1 ) = 0
∂t
∂~u1 ρ1 ∂~u1 ρ1 ∇P1
+ + ~u1 · ∇~u1 + ~u1 · ∇~u1 = −
∂t ρ0 ∂t ρ0 ρ0
The latter follows from first multiplying the momentum equations with (ρ0 + ρ1 )/ρ0 .
Note that we don’t need to consider the energy equation; this is because (i) we have
assumed that conduction is negligble, and (ii) the disturbance is adiabatic (meaning
dQ = 0, and there is thus no heating or cooling).

Next we linearize these equations, which means we use that the perturbed values
are all small such that terms that contain products of two or more of these quantities
are always negligible compared to those that contain only one such quantity. Hence,
the above equations reduce to
∂ρ1
+ ρ0 ∇~u1 = 0
∂t
∂~u1 ∇P1
+ = 0
∂t ρ0

86
These equations describe the evolution of perturbations in an inviscid and uniform
fluid. As always, these equations need an additional equation for closure. As men-
tioned above, we don’t need the energy equation: instead, we can use that the
flow is adiabatic, which implies that P ∝ ργ .

Using Taylor series expansion, we then have that


 
∂P
P (ρ0 + ρ1 ) = P (ρ0 ) + ρ1 + O(ρ21 )
∂ρ 0

where we have used (∂P/∂ρ)0 as shorthand for the partial derivative of P (ρ) at
ρ = ρ0 . And since the flow is isentropic, we have that the partial derivative is for
constant entropy. Using that P (ρ0 ) = P0 and P (ρ0 + ρ1 ) = P0 + P1 , we find that,
when linearized,  
∂P
P1 = ρ1
∂ρ 0
Note that P1 6= P (ρ1 ); rather P1 is the perturbation in pressure associated with the
perturbation ρ1 in the density.

Substitution in the fluid equations of our perturbed quantities yields


∂ρ1
+ ρ0 ∇~u1 = 0
∂t 
∂~u1 ∂P ∇ρ1
+ = 0
∂t ∂ρ 0 ρ0

Taking the partial time derivative of the above continuity equation, and using that
∂ρ0 /∂t = 0, gives
∂ 2 ρ1 ∂~u1
2
+ ρ0 ∇ · =0
∂t ∂t
Substituting the above momentum equation, and realizing that (∂P/∂ρ)0 is a
constant, then yields
 
∂ 2 ρ1 ∂P
− ∇2 ρ1 = 0
∂t2 ∂ρ 0

which we recognize as a wave equation, whose solution is a plane wave:


~
ρ1 ∝ ei(k·~x−ωt)

87
with ~k the wavevector, k = |~k| = 2π/λ the wavenumber, λ the wavelength,
ω = 2πν the angular frequency, and ν the frequency.

To gain some insight, consider the 1D case: ρ1 ∝ ei(kx−ωt) ∝ eik(x−vp t) , where we have
defined the phase velocity vp ≡ ω/k. This is the velocity with which the wave
pattern propagates through space. For our perturbation of a compressible fluid, this
phase velocity is called the sound speed, cs . Substituting the solution ρ1 ∝ ei(kx−ωt)
into the wave equation, we see that
s 
ω ∂P
cs = =
k ∂ρ s

where we have made it explicit that the flow is assumed to be isentropic. Note that
the partial derivative is for the unperturbed medium. This sound speed is sometimes
called the adiabatic speed of sound, to emphasize that it relies on the assumption
of an adiabatic perturbation. If the fluid is an ideal gas, then
s
kB T
cs = γ
µ mp

which shows that the adiabatic sound speed of an ideal fluid increases with temper-
ature.

We can repeat the above derivation by relaxing the assumption of isentropic flow,
and assuming instead that (more generally) the flow is polytropic. In that case,
P ∝ ρΓ , with Γ the polytropic index (Note: a polytropic EoS is an example of a
barotropic EoS). The only thing that changes is that now the sound speed becomes
s s
∂P P
cs = = Γ
∂ρ ρ

which shows that the sound speed is larger for a stiffer EoS (i.e., a larger value of Γ).

Note also that, for our barotropic fluid, the sound speed is independent of ω. This
implies that all waves move equally fast; the shape of a wave packet is preserved

88
as it moves. We say that an ideal (inviscid) fluid with a barotropic EoS is a non-
dispersive medium.

To gain further insight, let us look once more at the (1D) solution for our perturba-
tion:

ρ1 ∝ ei(kx−ωt) ∝ eikx e−iωt

Recalling Euler’s formula (eiθ = cos θ + i sin θ), we see that:

• The eikx part describes a periodic, spatial oscillation with wavelength λ = 2π/k.

• The e−iωt part describes the time evolution:

– If ω is real, then the solution describes a sound wave which propagates


through space with a sound speed cs .
– If ω is imaginary then the perturation is either exponentially growing
(‘unstable’) or decaying (‘damped’) with time.

We will return to this in Chapter 15, when we discuss the Jeans stability criterion.

As discussed above, acoustic waves result from disturbances in a compressible fluid.


These disturbances may arise from objects being moved through the fluid. However,
sound waves can also be sourced by fluid motions themselves. A familiar example
is the noise from jet-engines; the noise emenates from the turbulent wake created
by engines. In astrophysics, turbulence will also typically create sound waves. In
general these sound waves will not have an important impact on the physics. A
potential exception is the heating of the ICM by sound waves created by turbulent
wakes created by AGN feedback.

89
CHAPTER 12

Shocks

When discussing sound waves in the previous chapter, we considered small (linear)
perturbations. In this Chapter we consider the case in which the perturbations are
large (non-linear). Typically, a large disturbance results in an abrupt discontinuity
in the fluid, called a shock. Note: not all discontinuities are shocks, but all shocks
are discontinuities.

Consider a polytropic EoS:


 Γ
ρ
P = P0
ρ0
The sound speed is given by
 1/2 s  (Γ−1)/2
∂P P ρ
cs = = Γ = cs,0
∂ρ ρ ρ0
If Γ = 1, i.e., the EoS is isothermal, then the sound speed is a constant, independent
of density of pressure. However, if Γ 6= 1, then the sound speed varies with the local
density. An important example, often encountered in (astro)physics is the adiabatic
EoS, for which Γ = γ (γ = 5/3 for a mono-atomic gas). In that case we have that cs
increases with density (and pressure, and temperature).

In our discussion of sound waves (Chapter 13), we used perturbation theory, in


which we neglected the ~u1 · ∇~u1 term. However, when the perturbations are not
small, this term is no longer negligble, and causes non-linearities to develop. The
most important of those, is the fact that the sound speed itself varies with density (as
we have seen above). This implies that the wave-form of the acoustic wave changes
with time; the wave-crest is moving faster than the wave-trough, causing an overall
steepening of the wave-form. This steepening continues until the wave-crest tries
to overtake the wave-trough, which is not allowed, giving rise to a shock front (see
Fig. 13).

90
Figure 13: The steepening of a sound wave into a shock due to non-linearity (i.e.,
sound speed depends on density).

91
Mach Number: if v is the flow speed of the fluid, and cs is the sound speed, then
the Mach number of the flow is defined as
v
M=
cs

Note: simply accelerating a flow to supersonic speeds does not necessarily generate
a shock. Shocks only arise when an obstruction in the flow causes a deceleration of
fluid moving at supersonic speeds. The reason is that disturbances cannot propagate
upstream, so that the flow cannot ‘adjust itself’ to the obstacle because there is no
way of propagating a signal (which always goes at the sound speed) in the upstream
direction. Consequently, the flow remains undisturbed until it hits the obstacle,
resulting in a discontinuous change in flow properties; a shock.

Structure of a Shock: Fig. 14 shows the structure of a planar shock. The shock
has a finite, non-zero width (typically a few mean-free paths of the fluid particles),
and separates the ‘up-stream’, pre-shocked gas, from the ‘down-stream’, shocked gas.

For reasons that will become clear in what follows, it is useful to split the downstream
region in two sub-regions; one in which the fluid is out of thermal equilibrium, with
net cooling L > 0, and, further away from the shock, a region where the downstream
gas is (once again) in thermal equilibrium (i.e., L = 0). If the transition between
these two sub-regions falls well outside the shock (i.e., if x3 ≫ x2 ) the shock is said
to be adiabatic. In that case, we can derive a relation between the upstream (pre-
shocked) properties (ρ1 , P1 , T1 , u1) and the downstream (post-shocked) properties
(ρ2 , P2 , T2 , u2 ); these relations are called the Rankine-Hugoniot jump conditions.
Linking the properties in region three (ρ3 , P3 , T3 , u3) to those in the pre-shocked gas
is in general not possible, except in the case where T3 = T1 . In this case one may
consider the shock to be isothermal.

Rankine-Hugoniot jump conditions: We now derive the relations between the


up- and down-stream quantities, under the assumption that the shock is adiabatic.
Consider a rectangular volume V that encloses part of the shock; it has a thickness
dx > (x2 − x1 ) and is centered in the x-direction on the middle of shock. At fixed
x the volume is bounded by an area A. If we ignore variations in ρ and ~u in the y-

92
ρ1 P1 s ρ2 P2 ρ3 P3
h v2 T2
v1 T1 v3 T3
o
c
L=0 k L>0 L=0
x1 x2 x3
Figure 14: Structure of a planar shock.

and z-directions, the continuity equation becomes


∂ρ ∂
+ (ρ ux ) = 0
∂t ∂x
If we integrate this equation over our volume V we obtain
Z Z Z Z Z Z
∂ρ ∂
dx dy dz + (ρux ) dx dy dz = 0
∂t ∂x
Z Z
∂ ∂
⇔ ρ dx dy dz + A (ρux ) dx = 0
∂t ∂x
Z
∂M
⇔ + d(ρux ) = 0
∂t
Since there is no mass accumulation in the shock, and mass does not dissapear in
the shock, we have that
ρux |+dx/2 = ρux |−dx/2

In terms of the upstream (index 1) and downstream (index 2) quantities:

ρ1 u1 = ρ2 u2
This equation describes mass conservation across a shock.

The momentum equation in the x-direction, ignoring viscosity, is given by

93
∂ ∂ ∂Φ
(ρ ux ) = − (ρ ux ux + P ) − ρ
∂t ∂x ∂x
Integrating this equation over V and ignoring any gradient in Φ across the shock, we
obtain
ρ1 u21 + P1 = ρ2 u22 + P2
This equation describes how the shock converts ram pressure into thermal
pressure.

Finally, applying the same to the energy equation under the assumption that the
shock is adiabatic (i.e., dQ/dt = 0), one finds that (E + P )u has to be the same on
both sides of the shock, i.e.,
 
1 2 P
u +Φ+ε+ ρ u = constant
2 ρ
We have already seen that ρ u is constant. Hence, if we once more ignore gradients
in Φ across the shock, we obtain that

1 2 1
u1 + ε1 + P1 /ρ1 = u22 + ε2 + P2 /ρ2
2 2
This equation describes how the shock converts kinetic energy into enthalpy.
Qualitatively, a shock converts an ordered flow upstream into a disordered (hot) flow
downstream.

The three equations in the rectangular boxes are known as the Rankine-Hugoniot
(RH) jump conditions for an adiabatic shock. Using straightforward but
tedious algebra, these RH jump conditions can be written in a more useful form
using the Mach number M1 of the upstream gas:
  −1
ρ2 u1 1 γ −1 1
= = + 1−
ρ1 u2 M21 γ + 1 M21
P2 2γ γ−1
= M21 −
P1 γ+1 γ+1
   
T2 P2 ρ2 γ −1 2 2 1 4γ γ −1
= = γM1 − + −
T1 P1 ρ1 γ+1 γ+1 M21 γ−1 γ+1

94
Here we have used that for an ideal gas
kB T
P = (γ − 1) ρ ε = ρ
µ mp

Given that M1 > 1, we see that ρ2 > ρ1 (shocks compress), u2 < u1 (shocks
decelerate), P2 > P1 (shocks increase pressure), and T2 > T1 (shocks heat).
The latter may seem surprising, given that the shock is considered to be adiabatic:
although the process has been adiabatic, in that dQ/dt = 0, the gas has changed its
adiabat; its entropy has increased as a consequence of the shock converting kinetic
energy into thermal, internal energy. In general, in the presence of viscosity, a
change that is adiabatic does not imply that the states before and after are simply
linked by the relation P = K ργ , with K some constant. Shocks are always viscous,
which causes K to change across the shock, such that the entropy increases; it is this
aspect of the shock that causes irreversibility, thus defining an ”arrow of time”.

Back to the RH jump conditions: in the limit M1 ≫ 1 we have that


γ+1
ρ2 = ρ1 = 4 ρ1
γ−1

where we have used that γ = 5/3 for a monoatomic gas. Thus, with an adia-
batic shock you can achieve a maximum compression in density of a factor four!
Physically, the reason why there is a maximal compression is that the pressure and
temperature of the downstream fluid diverge as M21 . This huge increase in down-
stream pressure inhibits the amount of compression of the downstream gas. However,
this is only true under the assumption that the shock is adiabtic. The downstream,
post-shocked gas is out of thermal equilibrium, and in general will be cooling (i.e.,
L > 0). At a certain distance past the shock (i.e., when x = x3 in Fig. 14), the
fluid will re-establish thermal equilibrium (i.e., L = 0). In some special cases, one
can obtain the properties of the fluid in the new equilibrium state; one such case is
the example of an isothermal shock, for which the downstream gas has the same
temperature as the upstream gas (i.e., T3 = T1 ).

In the case of an isothermal shock, the first two Rankine-Hugoniot jump con-

95
ditions are still valid, i.e.,

ρ1 u1 = ρ3 u3
ρ1 u21 + P1 = ρ3 u23 + P3

However, the third condition, which derives from the energy equation, is no longer
valid. After all, in deriving that one we had assumed that the shock was adiabatic.
In the case of an isothermal shock we have to replace the third RH jump condition
with T1 = T3 . The latter implies that c2s = P3 /ρ3 = P1 /ρ1 , and allows us to rewrite
the second RH condition as

ρ1 (u21 + c2s ) = ρ3 (u23 + c2s )


⇔ u21 − ρρ13 u23 = ρρ13 c2s − c2s
⇔ u21 − u1 u3 = ( uu13 − 1) c2s
⇔ u1 u3 (u1 − u3 ) = (u1 − u3 ) c2s
⇔ c2s = u1 u3

Here the second step follows from using the first RH jump condition. If we now
substitute this result back into the first RH jump condition we obtain that
 2
ρ3 u1 u1
= = = M21
ρ1 u3 cs

Hence, in the case of isothermal shock (or an adiabatic shock, but sufficiently far
behind the shock in the downstream fluid), we have that there is no restriction to
how much compression the shock can achieve; depending on the Mach number of the
shock, the compression can be huge.

96
Figure 15: An actual example of a supernova blastwave. The red colors show the
optical light emitted by the supernova ejecta, while the green colors indicate X-ray
emission coming from the hot bubble of gas that has been shock-heated when the
blast-wave ran over it.

Supernova Blastwave: An important example of a shock in astrophysics are su-


pernova blastwaves. When a supernova explodes, it blasts a shell of matter (the
‘ejecta’) at high (highly supersonic) speed into the surrounding medium. The ki-
netic energy of this shell material is roughly ESN = 1051 erg. This is roughly 100
times larger than the amount of energy emitted in radiation by the supernova explo-
sion (which is what we ‘see’). For comparison, the entire Milky Way has a luminosity
of ∼ 1010 L⊙ ≃ 4 × 1043 ergs−1 , which amounts to an energy emitted by stars over an
entire year that is of the order of 1.5 × 1051 erg. Hence, the kinetic energy released
by a single SN is larger than the energy radiated by stars, by the entire galaxy, in
an entire year!
The mass of the ejecta is of the order of 1 Solar mass, which implies (using that
ESN = 12 Mej vej2 ), that the ejecta have a velocity of ∼ 10, 000 km s−1 !! Initially, this
shell material has a mass that is much larger than the mass of the surroundings swept
up by the shock, and to lowest order the shell undergoes free expansion. This phase
is therefore called the free-expansion phase. As the shock moves out, it engulves
more and more interstellar material, which is heated (and compressed) by the shock.
Hence, the interior of the shell (=shock) is a super-hot bubble of over-pressurized

97
gas, which ‘pushes’ the shock outwards. As more and more material is swept-up, and
accelerated outwards, the mass of the shell increases, which causes the velocity of the
shell to decelerate. At the early stages, the cooling of the hot bubble is negligble, and
the blastwave is said to be in the adiabatic phase, also known as the Sedov-Taylor
phase. At some point, though, the hot bubble starts to cool, radiating away the
kinetic energy of the supernova, and lowering the interior pressure up to the point
that it no longer pushes the shell outwards. This is called the radiative phase.
From this point on, the shell expands purely by its inertia, being slowed down by the
work it does against the surrounding material. This phase is called the snow-plow
phase. Ultimately, the velocity of the shell becomes comparable to the sound speed
of the surrounding material, after which it continues to move outward as a sound
wave, slowly dissipating into the surroundings.

During the adiabatic phase, we can use a simple dimensional analysis to solve for
the evolution of the shock radius, rsh , with time. Since the only physical parameters
that can determine rsh in this adiabatic phase are time, t, the initial energy of the
SN explosion, ε0 , and the density of the surrounding medium, ρ0 , we have that

rsh = f (t, ε0 , ρ0 ) = Atη εα0 ρβ0

It is easy to check that there is only one set of values for η, α and β for which the
product on the right has the dimensions of length (which is the dimension of rsh .
This solution has η = 2/5, α = 1/5 and β = −1/5, such that
 1/5
ε
rsh = A t2/5
ρ0

and thus  1/5


drsh 2A ε
vsh = = t−3/5
dt 5 ρ0
which shows that indeed the shock decelerates as it moves outwards.

98
CHAPTER 13

Fluid Instabilities

In this Chapter we discuss the following instabilities:


• convective instability (Schwarzschild criterion)
• interface instabilities (Rayleigh-Taylor & Kelvin-Helmholtz)
• gravitational instability (Jeans criterion)
• thermal instability (Field criterion)

Convective Instability: In astrophysics we often need to consider fluids heated


from ”below” (e.g., stars, Earth’s atmosphere where Sun heats surface, etc.)3 . This
results in a temperature gradient: hot at the base, colder further ”up”. Since warmer
fluids are more buoyant (‘lighter’), they like to be further up than colder (‘heavier’)
fluids. The question we need to address is under what conditions this adverse tem-
perature gradient becomes unstable, developing ”overturning” motions known as
thermal convection.

Consider a blob with density ρb and pressure Pb embedded in an ambient medium of


density ρ and pressure P . Suppose the blob is displaced
by a small distance δz upward. After the displace-
ment the blob will have conditions (ρ∗b , Pb∗ ) and its
new ambient medium is characterized by (ρ′ , P ′ ),
where
dρ dP
ρ′ = ρ + δz P′ = P + δz
dz dz
Initially the blob is assumed to be in mechani-
cal and thermal equilibrium with its ambient
medium, so that ρb = ρ and Pb = P . After the
displacement the blob needs to re-establish a new
mechanical and thermal equilibrium. In general,
the time scale on which it re-establishes mechan-
ical (pressure) equilibrium is the sound crossing
3
Here and in what follows, ‘up’ refers to the direction opposite to that of gravity.

99
time, τs , while re-establishing thermal equilibrium
proceeds much slower, on the conduction time, τc . Given that τs ≪ τc we can assume
that Pb∗ = P ′ , and treat the displacement as adiabatic. The latter implies that the
process can be described by an adiabatic EoS: P ∝ ργ . Hence, we have that
 1/γ  1/γ  1/γ
Pb∗ P′ 1 dP
ρ∗b = ρb = ρb = ρb 1+ δz
Pb P P dz

In the limit of small displacements δz, we can use Taylor series expansion to show
that, to first order,
ρ dP
ρ∗b = ρ + δz
γ P dz
where we have used that initially ρb = ρ, and that the Taylor series expansion,
f (x) ≃ f (0)+f ′(0)x+ 21 f ′′ (0)x2 +..., of f (x) = [1+x]1/γ is given by f (x) ≃ 1+ γ1 x+....
Suppose we have a stratified medium in which dρ/dz < 0 and dP/dz < 0. In that
case, if ρ∗b > ρ′ the blob will be heavier than its surrounding and it will sink back to its
original position; the system is stable to convection. If, on the other hand, ρ∗b < ρ′
then the displacement has made the blob more buoyant, resulting in instability.
Hence, using that ρ′ = ρ + (dρ/dz) δz we see that stability requires that

dρ ρ dP
<
dz γ P dz
This is called the Schwarzschild criterion for convective stability.

It is often convenient to rewrite this criterion in a form that contains the temperature.
Using that
µ mp
ρ = ρ(P, T ) = P
kB T
it is straightforward to show that
dρ ρ dP ρ dT
= −
dz P dz T dz
Substitution in ρ′ = ρ + (dρ/dz) δz then yields that
 
∗ ′ 1 ρ dP ρ dT
ρb − ρ = −(1 − ) + δz
γ P dz T dz

100
Since stability requires that ρ∗b − ρ′ > 0, and using that δz > 0, dP/dz < 0 and
dT /dz < 0 we can rewrite the above Schwarzschild criterion for stability as
 
dT 1 T dP
< 1−
dz γ P dz

This shows that if the temperature gradient becomes too large the system becomes
convectively unstable: blobs will rise up until they start to loose their thermal en-
ergy to the ambient medium, resulting in convective energy transport that tries to
“overturn” the hot (high entropy) and cold (low entropy) material. In fact, without
any proof we mention that in terms of the specific entropy, s, one can also write
the Schwarzschild criterion for convective stability as ds/dz > 0.

To summarize, the Schwarzschild criterion for convective stability is given by


either of the following three expressions:

 
dT 1 T dP
< 1−
dz γ P dz

dρ ρ dP
<
dz γ P dz

ds
>0
dz

Rayleigh-Taylor Instability: The Rayleigh—Taylor (RT) instability is an insta-


bility of an interface between two fluids of different densities that occurs when one
of the fluids is accelerated into the other. Examples include supernova explosions
in which expanding core gas is accelerated into denser shell gas and the common
terrestrial example of a denser fluid such as water suspended above a lighter fluid
such as oil in the Earth’s gravitational field.

It is easy to see where the RT instability comes from. Consider a fluid of density
ρ2 sitting on top of a fluid of density ρ1 < ρ2 in a gravitational field that is point-
ing in the downward direction. Consider a small perturbation in which the initially
horizontal interface takes on a small amplitude, sinusoidal deformation. Since this

101
Figure 16: Example of Rayleigh-Taylor instability in a hydro-dynamical simulation.

implies moving a certain volume of denser material down, and an equally large vol-
ume of the lighter material up, it is immediately clear that the potential energy of
this ‘perturbed’ configuration is lower than that of the initial state, and therefore
energetically favorable. Simply put, the initial configuration is unstable to small
deformations of the interface.

Stability analysis (i.e., perturbation analysis of the fluid equations) shows that the
dispersion relation corresponding to the RT instability is given by
r
g ρ2 − ρ1
ω = ±i k
k ρ2 + ρ1
where g is the gravitational acceleration, and the factor (ρ2 − ρ1 )/(ρ2 + ρ1 ) is called
the Atwood number. Since the wavenumber of the perturbation k > 0 we see
that ω is imaginary, which implies that the perturbations will grow exponentially
(i.e., the system is unstable). If ρ1 > ρ2 though, ω is real, and the system is stable
(perturbations to the interface propagate as waves).

Kelvin-Helmholtz Instability: the Kelvin-Helmholtz (KH) instability is an in-


terface instability that arises when two fluids with different densities have a velocity
difference across their interface. Similar to the RT instability, the KH instability
manifests itself as a small wavy pattern in the interface which develops into turbu-
lence and which causes mixing. Examples where KH instability plays a role are wind
blowing over water, (astrophysical) jets, the cloud bands on Jupiter (in particular
the famous red spot), and clouds of denser gas falling through the hot, low density
intra-cluster medium (ICM).

102
Figure 17: Illustration of onset of Kelvin-Helmholtz instability

Stability analysis (i.e., perturbation analysis of the fluid equations) shows that the
the dispersion relation corresponding to the KH instability is given by

ω (ρ1 u1 + ρ2 u2 ) ± i (u1 − u2 ) (ρ1 ρ2 )1/2


=
k ρ1 + ρ2
Note that this dispersion relation has both real and imaginary parts, given by

ωR (ρ1 u1 + ρ2 u2 )
=
k ρ1 + ρ2
and
ωI (ρ1 ρ2 )1/2
= (u1 − u2 )
k ρ1 + ρ2
Since the imaginary part is non-zero, except for u1 = u2 , we we have that, in principle,
any velocity difference across an interface is KH unstable. In practice, surface
tension can stabilize the short wavelength modes so that typically KH instability
kicks in above some velocity treshold.

As an example, consider a cold cloud of radius Rc falling into a cluster of galax-


ies. The latter contains a hot intra-cluster medium (ICM), and as the cloud moves
through this hot ICM, KH instabilities can develop on its surface. If the cloud started
out at a large distance from the cluster with zero velocity, than at infall it has a ve-
locity v ∼ vesc ∼ cs,h , where the latter is the sound speed of the hot ICM, assumed
to be in hydrostatic equilibrium. Defining the cloud’s overdensity δ = ρc /ρh − 1, we
can write the (imaginary part of the) dispersion relation as

ρh (ρc /ρh )1/2 (δ + 1)1/2


ω= cs,h k = cs,h k
ρh [1 + (ρc /ρh )] δ+2

103
The mode that will destroy the cloud has k ∼ 1/Rc , so that the time-scale for cloud
destruction is
1 Rc δ + 2
τKH ≃ ≃
ω cs,h (δ + 1)1/2
Assuming pressure equilibrium between cloud and ICM, and adopting the EoS of an
ideal gas, implies that ρh Th = ρc Tc , so that
1/2 1/2
cs,h T ρc
= h1/2 = 1/2 = (δ + 1)1/2
cs,c Tc ρh
Hence, one finds that the Kelvin-Helmholtz time for cloud destruction is
1 Rc δ + 2
τKH ≃ ≃
ω cs,c δ + 1

Note that τKH ∼ ζ(Rc/cs,c ) = ζτs , with ζ = 1(2) for δ ≫ 1(≪ 1). Hence, the Kelvin-
Helmholtz instability will typically destroy clouds falling into a hot ”atmosphere”
on a time scale between one and two sound crossing times, τs , of the cloud. Note,
though, that magnetic fields and/or radiative cooling at the interface may stabilize
the clouds.

Gravitational Instability: In our discussion of sound waves we used perturbation


analysis to derive a dispersion relation ω 2 = k 2 c2s . In deriving that equation we
ignored gravity by setting ∇Φ = 0 (see Chapter 13). If you do not ignore gravity,
then you add one more perturbed quantity; Φ = Φ0 + Φ1 and one more equation,
namely the Poisson equation ∇2 Φ = 4πGρ.
It is not difficult to show that this results in a modified dispersion relation:

ω 2 = k 2 c2s − 4πGρ0 = c2s k 2 − kJ2

where we have introduced the Jeans wavenumber



4πGρ0
kJ =
cs
to which we can also associate a Jeans length
r
2π π
λJ ≡ = cs
kJ Gρ0

104
and a Jeans mass  3
4 λJ π
MJ = πρ0 = ρ0 λ3J
3 2 6
From the dispersion relation one immediately sees that the system is unstable (i.e.,
ω is imaginary) if k < kJ (or, equivalently, λ > λJ or M > MJ ). This is called the
Jeans criterion for gravitational instability. It expresses when pressure forces
(which try to disperse matter) are no longer able to overcome gravity (which tries to
make matter collapse), resulting in exponential gravitational collapse on a time scale
r

τff =
32 G ρ

known as the free-fall time for gravitational collapse.

The Jeans stability criterion is of utmost importance in astrophysics. It is used to


describes the formation of galaxies and large scale structure in an expanding space-
time (in this case the growth-rate is not exponential, but only power-law), to describe
the formation of stars in molecular clouds within galaxies, and it may even play an
important role in the formation of planets in protoplanetary disks.

In deriving the Jeans Stability criterion you will encounter a somewhat puzzling issue.
Consider the Poisson equation for the unperturbed medium (which has density ρ0
and gravitational potential Φ0 ):

∇2 Φ0 = 4πGρ0

Since the initial, unperturbed medium is supposed to be homogeneous there can be


no gravitational force; hence ∇Φ0 = 0 everywhere. The above Poisson equation
then implies that ρ0 = 0. In other words, an unperturbed, homogeneous density
field of non-zero density does not seem to exist. Sir James Jeans ‘ignored’ this
‘nuisance’ in his derivation, which has since become known as the Jeans swindle.
The problem arises because Newtonian physics is not equipped to deal with systems
of infinite extent (a requirement for a perfectly homogeneous density distribution).
See Kiessling (1999; arXiv:9910247) for a detailed discussion, including an elegant
demonstration that the Jeans swindle is actually vindicated!

105
Figure 18: The locus of ther-
mal equilibrium (L = 0) in
the (ρ, T ) plane, illustrating the
principle of thermal instability.
The dashed line indicates a line
of constant pressure.

Thermal Instability: Let L = L(ρ, T ) = C − H be the net cooling rate. If L = 0


the system is said to be in thermal equilibrium (TE), while L > 0 and L < 0
correspond to cooling and heating, respectively.

The condition L(ρ, T ) = 0 corresponds to a curve in the (ρ, T )-plane with a shape
similar to that shown in Fig. 18. It has flat parts at T ∼ 106 K, at T ∼ 104K, at
T ∼ 10 − 100K. This can be understood from simple atomic physics (see for example
§ 8.5.1 of Mo, van den Bosch & White, 2010). Above the TE curve we have that
L > 0 (net cooling), while below it L < 0 (net heating). The dotted curve indicates
a line of constant pressure (T ∝ ρ−1 ). Consider a blob in thermal and mechanical
(pressure) equilibrium with its ambient medium, and with a pressure indicated by
the dashed line. There are five possible solutions for the density and temperature of
the blob, two of which are indicated by P1 and P2 ; here confusingly the P refers to
‘point’ rather than ‘pressure’. Suppose I have a blob located at point P2 . If I heat
the blob, displacing it from TE along the constant pressure curve (i.e., the blob is
assumed small enough that the sound crossing time, on which the blob re-established
mechanical equilibrium, is short). The blob now finds itself in the region where L > 0
(i.e, net cooling), so that it will cool back to its original location on the TE-curve;
the blob is stable. For similar reasons, it is easy to see that a blob located at point
P1 is unstable. This instability is called thermal instability, and it explains
why the ISM is a three-phase medium, with gas of three different temperatures
(T ∼ 106 K, 104 K, and ∼ 10 − 100 K) coexisting in pressure equilibrium. Gas at any
other temperature but in pressure equilibrium is thermally unstable.

106
It is easy to see that the requirement for thermal instability translates into
 
∂L
<0
∂T P

which is known as the Field criterion for thermal instability (after astrophysicist
George B. Field).

Fragmentation and Shattering: Consider the Jeans criterion, expressing a bal-


ance between gravity and pressure. Using that the Jeans mass MJ ∝ ρ λ3J and that
λJ ∝ ρ−1/2 cs , we see that
MJ ∝ ρ−1/2 T 3/2
where we have used that cs ∝ T 1/2 . Now consider a polytropic equation of state,
which has P ∝ ρΓ , with Γ the polytropic index. Assuming an ideal gas, such that
kB T
P = ρ
µ mp

we thus see that a polytropic ideal gas must have that T ∝ ρΓ−1 . Substituting that
in the expression for the Jeans mass, we obtain that
3 3 4
MJ ∝ ρ 2 Γ−2 = ρ 2 (Γ− 3 )

Thus, we see that for Γ > 4/3 the Jeans mass will increase with increasing density,
while the opposite is true for Γ < 4/3. Now consider a system that is (initially)
larger than the Jeans mass. Since pressure can no longer support it against its own
gravity, the system will start to collapse, which increases the density. If Γ < 4/3,
the Jeans mass will becomes smaller as a consequence of the collapse, and now small
subregions of the system will find themselves having a mass larger than the Jeans
mass ⇒ the system will start to fragment.

If the collapse is adiabatic (i.e., we can ignore cooling), then Γ = γ = 5/3 > 4/3 and
there will be no fragmentation. However, if cooling is very efficient, such that while
the cloud collapses it maintains the same temperature, the EoS is now isothermal,
which implies that Γ = 1 < 4/3: the cloud will fragment into smaller collapsing
clouds. Fragmentation is believed to underly the formation of star clusters.

107
A very similar process operates related to the thermal instability. In the discussion
of the Field criterion we had made the assumption “the blob is assumed small
enough that the sound crossing time, on which the blob re-established mechanical
equilibrium, is short”. Here ‘short’ means compared to the cooling time of the cloud.
Let’s define the cooling length lcool ≡ cs τcool , where cs is the cloud’s sound speed and
τcool is the cooling time (the time scale on which it radiates away most of its internal
energy). The above assumption thus implies that the size of the cloud, lcloud ≪
lcool . As a consequence, whenever the cloud cools somewhat, it can immediately
re-establish pressure equilibrium with its surrounding (i.e., the sound crossing time,
τs = lcloud /cs is much smaller than the cooling time τcool = lcool /cs ).

Now consider a case in which lcloud ≫ lcool (i.e., τcool ≪ τs ). As the cloud cools,
it cannot maintain pressure equilibrium with its surroundings; it takes too long for
mechanical equilibrium to be established over the entire cloud. What happens is that
smaller subregions, of order the size lcool , will fragment. The smaller fragments will
be able to maintain pressure equilibrium with their surroundings. But as the small
cloudlets cool further, the cooling length lcool shrinks. To see this, realize that when T
drops this lowers the sound speed and decreases the cooling time; after all, we are in
the regime of thermal instability, so (∂L/∂T )P < 0. As a consequence, lcool = cs τcool
drops as well. So the small cloudlett soon finds itself larger than the cooling length,
and it in turn will fragment. This process of shattering continues until the cooling
time becomes sufficiently long and the cloudletts are no longer thermally unstable
(see McCourt et al., 2018, MNRAS, 473, 5407 for details).

This process of shattering is believed to play an important role in the inter-galactic


medium (IGM) in between galaxies, and the circum-galactic medium (CGM) in the
halos of galaxies.

108
Part II: Collisionless Dynamics

The following chapters give an elementary introduction into the rich topic of colli-
sionless dynamics. The main goal is to highlight how the lack of collisions among the
constituent particles give rise to a dynamics that differs remarkably from collisional
fluids. We also briefly discuss the theory or orbits, which are the building blocks of
collisionless systems, the Virial theorem, and the gravothermal catastrophe, which
is a consequence of the negative heat capacity of a gravitational system. Finally, we
briefly discuss interactions (‘collisions’) among collisionless systems.

Collisionless Dynamics is a rich topic, and one could easily devote an entire course
to it (for example the Yale Graduate Course ‘ASTR 518; Galactic Dynamics’). The
following chapters therefore only scratch the surface of this rich topic. Readers who
want to get more indepth information are referred to the following excellent text-
books
- Galactic Dynamics by J. Binney & S. Tremaine
- Galactic Nuclei by D. Merritt
- Galaxy Formation and Evolution by H.J. Mo, F. van den Bosch & S. White

109
CHAPTER 14

Potential Theory & The Virial Theorem

Gravity in Astrophysical Fluids: Many of the fluids encountered in astrophysics


are self-gravitating, which means that the gravitational force due to the fluid itself
exceeds the gravitational force from the external mass distribution. Arguably the
most important example of self-gravitating, astrophysical fluids are stars. But Cold
Dark Matter halos are also examples of self-gravitating fluids (albeit collisionless).
The interstellar medium (ISM) can and cannot be self-gravitating, depending on the
conditions. The intra-cluster medium (ICM) is generally not self-gravitating; rather
the gravitating potential is dominated by the dark matter.

Gravitational Potential: Gravity is a conservative force, which means that it can


be written as the gradient of a scalar field. Newton’s gravitational potential, Φ(~x),
is defined such that the gravitational force per unit mass
F~g = −∇Φ
Note that the absolute normalization of Φ has no physical relevance; only the gradi-
ents of Φ matter.

Consider a density distribution ρ(~x). What is the gravitational force F~g acting on
a particle of mass m at location ~x ? We can sum the small constributions δ F~g from
different regions ~x ′ ± d3~x ′ , given by

m δm(~x ′ ) ~x ′ − ~x ~x ′ − ~x
δ F~g (~x) = G ′ 2 ′
= Gm ′ 3
ρ(~x ′ )d3~x ′
|~x − ~x| |~x − ~x| |~x − ~x|
R
Adding up all the small contributions yields F~g (~x) = δ F~g (~x) ≡ m ~g (~x), where
Z
~x ′ − ~x
~g (~x) = G d3~x ′ ′ ρ(~x ′ )
|~x − ~x|3
is the gravitational field (i.e., the force per unit mass). Using that
 
~x ′ − ~x 1
= ∇x
|~x ′ − ~x|3 |~x ′ − ~x|

110
we can rewrite g(~x) as
Z   Z
3 ′ 1 ′ Gρ(~x ′ )
~g (~x) = G d ~x ∇x ρ(~x ) = ∇x d3~x ′ ≡ −∇x Φ
|~x ′ − ~x| |~x ′ − ~x|
where in the last step we have defined the gravitational potential
Z
ρ(~x ′ )
Φ(~x) = −G d3~x ′ ′
|~x − ~x|

It can be shown, that the above expression is equivalent to what is known as the
Poisson equation:

∇2 Φ = 4π G ρ
For a derivation, see Section 3.2 of Astrophysical Fluid Dynamics by Clarke &
Carswell, or Section 2.1 of Galactic Dynamics by Binney & Tremaine.

In general, it is extremely complicated to solve the Poisson equation for Φ(~x) given
ρ(~x) [see Chapter 2 of Galactic Dynamics by Binney & Tremaine for a detailed
discussion]. However, under certain symmetries, solutions to the Poisson equation
are fairly straightforward. In particular, under spherical symmetry the general
solution to the Poisson equation is
 Z r Z ∞ 
1 ′ ′2 ′ ′ ′ ′
Φ(r) = −4πG ρ(r ) r dr + ρ(r ) r dr
r 0 r

Note that the potential at r depends on the mass distribution outside of r. However,
if we now compute the gravitational force per unit mass

dΦ G M(r)
F~g (r) = − êr = − êr
dr r2
where Z r
M(r) ≡ 4π ρ(r ′ ) r ′2 dr
0
is the enclosed mass within r. This shows that the gravitational force does not
depend on the mass distribution outside of r.

111
Newton’s first theorem: a body that is inside a spherical shell of matter experi-
ences no net gravitational force from that shell. The equivalent in general relativity
is called Birkhoff’s theorem.

This is easily understood from the fact that the solid angles that extent from a
point inside a sphere to opposing directions have areas on the sphere that scale as r 2
(where r is the distance from the point to the sphere), while the gravitational force
per unit mass scales as r −2 . Hence, the gravitational forces from the two opposing
areas exactly cancel.

Circular velocity: the velocity of a particle or fluid element on a circular orbit.


For a spherical mass distribution
r r
dΦ G M(r)
Vcirc (r) = r =
dr r
In the case of an axisymmetric mass distribution, the circular velocity in the equa-
torial plane (z = 0, where z is one of the three cylindrical coordinates (R, φ, z)) is
given by
r r
dΦ G M(R)
Vcirc (R) = R 6=
dR R

Escape velocity: the velocity needed for a particle or fluid element to escape to
infinity. Since E = v 2 /2 + Φ(~x), and escape requires E > 0, the escape velocity is
p
Vesc (~x) = 2 |Φ(~x)|
independent of the symmetry (or lack thereof) of the mass distribution.

Since gas cannot be on self-intersecting orbits, gas in disk galaxies generally orbits
on circular orbits. The measured rotation velocities therefore reflect the circular
velocities, which can be used to infer the enclosed mass as a function of radius. This
method is generaly used to infer the presence of dark matter halos surrounding
disk galaxies.

112
Consider a gravitational system consisting of N particles (e.g., stars, fluid elements).
The total energy of the system is E = K + W , where

P
N
1
Total Kinetic Energy: K= 2
mi vi2
i=1
P
N P
G mi mj
Total Potential Energy: W = − 21 |~
ri −~
rj |
i=1 j6=i

The latter follows from the fact that gravitational binding energy between a pair
of masses is proportional to the product of their masses, and inversely proportional
to their separation. The factor 1/2 corrects for double counting the number of pairs.

Potential Energy in Continuum Limit: To infer an expression for the gravi-


tational potential energy in the continuum limit, it is useful to rewrite the above
expression as
N
1X
W = mi Φi
2 i=1
where
X G mj
Φi = −
j6=i
rij

where rij = |~ri − ~rj |. In the continuum limit this simply becomes
Z
1
W = ρ(~x) Φ(~x) d3~x
2
One can show (see e.g., Binney & Tremaine 2008) that this is equal to the trace of
the Chandrasekhar Potential Energy Tensor
Z
∂Φ 3
Wij ≡ − ρ(~x) xi d ~x
∂xj
In particular,
3
X Z
W = Tr(Wij ) = Wii = − ρ(~x) ~x · ∇Φ d3~x
i=1

113
which is another, equally valid, expression for the gravitational potential energy in
the continuum limit.

Virial Theorem: A stationary, gravitational system obeys

2K + W = 0

Actually, the correct virial equation is 2K + W + Σ = 0, where Σ is the surface pressure.


In many, but certainly not all, applications in astrophysics this term can be ignored. Many
textbooks don’t even mention the surface pressure term.

Combining the virial equation with the expression for the total energy, E = K +W ,
we see that for a system that obeys the virial theorem

E = −K = W/2

Example: Consider a cluster consisting of N galaxies. If the cluster is in virial


equilibrium then
N
X N
1 1 X X G mi mj
2 m vi2 − =0
i=1
2 2 i=1 rij
j6=i

If we assume, for simplicity, that all galaxies have equal mass then we can rewrite
this as
N N
1 X 2 G (Nm)2 1 X X 1
Nm vi − =0
N i=1 2 N 2 i=1 rij
j6=i

Using that M = N m and N(N − 1) ≃ N 2 for large N, this yields

2 hv 2 i
M=
G h1/ri
with

114
XX 1 N
1
h1/ri =
N(N − 1) i=1 j6=i rij

It is useful to define the gravitational radius rg such that

G M2
W =−
rg
Using the relations above, it is clear that rg = 2/h1/ri. We can now rewrite the
above equation for M in the form

rg hv 2 i
M=
G
Hence, one can infer the mass of our cluster of galaxies from its velocity dispersion
and its gravitation radius. In general, though, neither of these is observable, and one
uses instead
2
Reff hvlos i
M =α
G
where vlos is the line-of-sight velocity, Reff is some measure for the ‘effective’ radius
of the system in question, and α is a parameter of order unity that depends on the
radial distribution of the galaxies. Note that, under the assumption of isotropy,
2
hvlos i = hv 2 i/3 and one can also infer the mean reciprocal pair separation from the
projected pair separations; in other words under the assumption of isotropy one can
infer α, and thus use the above equation to compute the total, gravitational mass of
the cluster. This method was applied by Fritz Zwicky in 1933, who inferred that
the total dynamical mass in the Coma cluster is much larger than the sum of the
masses of its galaxies. This was the first observational evidence for dark matter,
although it took the astronomical community until the late 70’s to generally accept
this notion.

115
For a self-gravitating fluid
N
X 1 1 3
K= mi vi2 = N m hv 2 i = N kB T
i=1
2 2 2
where the last step follows from the kinetic theory of ideal gases of monoatomic
particles. In fact, we can use the above equation for any fluid (including a collisionless
one), if we interpret T as an effective temperature that measures the rms velocity
of the constituent particles. If the system is in virial equilibrium, then
3
E = −K = − N kB T
2
which, as we show next, has some important implications...

Heat Capacity: the amount of heat required to increase the temperature by one
degree Kelvin (or Celsius). For a self-gravitating fluid this is
dE 3
= − N kB
C≡
dT 2
which is negative! This implies that by losing energy, a gravitational system
gets hotter!! This is a very counter-intuitive result, that often leads to confusion and
wrong expectations. Below we give three examples of implications of the negative
heat capacity of gravitating systems,

Example 1: Drag on satellites Consider a satellite orbiting Earth. When it expe-


riences friction against the (outer) atmosphere, it loses energy. This causes the system
to become more strongly bound, and the orbital radius to shrink. Consequently, the
energy loss results in the gravitational potential energy, W , becoming more negative.
In order for the satellite to re-establish virial equilibrium (2K + W = 0), its kinetic
energy needs to increase. Hence, contrary to common intuition, friction causes
the satellite to speed up, as it moves to a lower orbit (where the circular velocity is
higher).

Example 2: Stellar Evolution A star is a gaseous, self-gravitating sphere that


radiates energy from its surface at a luminosity L. Unless this energy is replenished
(i.e., via some energy production mechanism in the star’s interior), the star will react
by shrinking (i.e., the energy loss implies an increase in binding energy, and thus a

116
potential energy that becomes more negative). In order for the star to remain in
virial equilibrium its kinetic energy, which is proportional to temperature, has to
increase; the star’s energy loss results in an increase of its temperature.

In the Sun, hydrogen burning produces energy that replenishes the energy loss from
the surface. As a consequence, the system is in equilibrium, and will not contract.
However, once the Sun has used up all its hydrogren, it will start to contract and heat
up, because of the negative heat capacity. This continues until the temperature in
the core becomes sufficiently high that helium can start to fuse into heavier elements,
and the Sun settles in a new equilibrium.

Example 3: Core Collapse a system with negative heat capacity in contact with
a heat bath is thermodynamically unstable. Consider a self-gravitating fluid of ‘tem-
perature’ T1 , which is in contact with a heat bath of temperature T2 . Suppose the
system is in thermal equilibrium, so that T1 = T2 . If, due to some small disturbance,
a small amount of heat is tranferred from the system to the heat bath, the negative
heat capacity implies that this results in T1 > T2 . Since heat always flows from hot
to cold, more heat will now flow from the system to the heat bath, further increasing
the temperature difference, and T1 will continue to rise without limit. This run-away
instability is called the gravothermal catastrophe. An example of this instability
is the core collapse of globular clusters: Suppose the formation of a gravitational
system results in the system having a declining velocity dispersion profile, σ 2 (r) (i.e.,
σ decreases with increasing radius). This implies that the central region is (dynami-
cally) hotter than the outskirts. IF heat can flow from the center to those outskirts,
the gravothermal catastrophe kicks in, and σ in the central regions will grow with-
out limits. Since σ 2 = GM(r)/r, the central mass therefore gets compressed into
a smaller and smaller region, while the outer regions expand. This is called core
collapse. Note that this does NOT lead to the formation of a supermassive black
hole, because regions at smaller r always shrink faster than regions at somewhat
larger r. In dark matter halos, and elliptical galaxies, the velocity dispersion profile
is often declining with radius. However, in those systems the two-body relaxation
time is soo long that there is basically no heat flow (which requires two-body in-
teractions). However, globular clusters, which consist of N ∼ 104 stars, and have
a crossing time of only tcross ∼ 5 × 106 yr, have a two-body relaxation time of only
∼ 5 × 108 yr. Hence, heat flow in globular clusters is not negligible, and they can
(and do) undergo core collapse. The collapse does not proceed indefinitely, because
of binaries (see Galactic Dynamics by Binney & Tremaine for more details).

117
CHAPTER 15

Collisionless Dynamics: CBE & Jeans Equations

In this chapter we consider collisionless fluids, such as galaxies and dark matter halos.
As discussed in previous chapters, their dynamics is governed by the Collisionless
Boltzmann equation (CBE)

df ∂f ∂f ∂Φ ∂f
= + vi − =0
dt ∂t ∂xi ∂xi ∂vi

By taking the velocity moment of the CBE (see Chapter 8), we obtain the Jeans
equations
∂ui ∂ui 1 ∂ σ̂ij ∂Φ
+ uj =− −
∂t ∂xj ρ ∂xj ∂xi

which are the equivalent of the Navier-Stokes equations (or Euler equations), but for
a collisionless fluid. The quantity σ̂ij in the above expression is the stress tensor,
defined as
σ̂ij = −ρ hwi wj i = −ρ(hvi vj i − hvi i hvj i)
In this chapter, we write a hat on top of the stress tensor, in order to distinguish it
from the velocity dispersion tensor given by

σ̂ij
σij2 = hvi vj i − hvi i hvj i = −
ρ

This notation may cause some confusion, but it is adapted here in order to be con-
sistent with the notation in standard textbooks on galactic dynamics. For the same
reason, in what follows we will write hvi i in stead of ui (also because ui was defined
as the velocity of a fluid element, but for a collisionless fluid the concept of a fluid
element is not defined).

As we have discussed in detail in Chapters 4 and 5, for a collisional fluid the stress
tensor is given by
σ̂ij = −ρσij2 = −P δij + τij

118
and therefore completely specified by two scalar quantities; the pressure P and the
shear viscosity µ (as always, we ignore bulk viscosity). Both P and µ are related to
ρ and T via constitutive equations, which allow for closure in the equations.
In the case of a collisionless fluid, though, no consistutive relations exist, and the
(symmetric) velocity dispersion tensor has 6 unknowns. As a consequence, the Jeans
equations do not form a closed set. Adding higher-order moment equations of the
CBE will yield more equations, but this also adds new, higher-order unknowns such
as hvi vj vk i, etc. As a consequence, the set of CBE moment equations never closes!

Note that σij2 is a local quantity; σij2 = σij2 (~x). At each point ~x it defines the
velocity ellipsoid; an ellipsoid whose principal axes are defined by the orthogonal
eigenvectors of σij2 with lengths that are proportional to the square roots of the
respective eigenvalues.

Since these eigenvalues are typically not the same, a collisionless fluid experiences
anisotropic pressure-like forces. In order to be able to close the set of Jeans equa-
tions, it is common to make certain assumptions about the symmetry of the fluid.
For example, a common assumption is that the fluid is isotropic, such that the
(local) velocity dispersion tensor is specified by a single quantity; the local velocity
dispersion σ 2 . Note, though, that if with this approach, a solution is found, the
solution may not correspond to a physical distribution function (DF) (i.e., in order
to be physical, f ≥ 0 everywhere). Thus, although any real DF obeys the Jeans
equations, not every solution to the Jeans equations corresponds to a physical DF!!!

As a worked out example, we now derive the Jeans equations under cylindrical
symmetry. We therefore write the Jeans equations in the cylindrical coordinate
system (R, φ, z). The first step is to write the CBE in cylindrical coordinates.

df ∂f ∂f ∂f ∂f ∂f ∂f ∂f
= + Ṙ + φ̇ + ż + v̇R + v̇φ + v̇z
dt ∂t ∂R ∂φ ∂z ∂vR ∂vφ ∂vz

Recall from vector calculus (see Appendices A and D) that

~v = Ṙ~eR + Rφ̇~eφ + ż~ez = vR~eR + vφ~eφ + vz~ez

from which we obtain the acceleration vector


d~v
~a = = R̈~eR + Ṙ~e˙ R + Ṙφ̇~eφ + Rφ̈~eφ + Rφ̇~e˙ φ + z̈~ez + ż~e˙ z
dt
119
Using that ~e˙ R = φ̇~eφ , ~e˙ φ = −φ̇~eR , and ~e˙ z = 0 we have that
h i h i
2
~a = R̈ − Rφ̇ ~eR + 2Ṙφ̇ + Rφ̈ ~eφ + z̈~ez

Next we use that

vR = Ṙ ⇒ v̇R = R̈
vφ = Rφ̇ ⇒ v̇φ = Ṙφ̇ + Rφ̈
vz = ż ⇒ v̇z = z̈

to write the acceleration vector as


  hv v i
vφ2 R φ
~a = v̇R − ~eR + + v̇φ ~eφ + v̇z ~ez
R R

Newton’s equation of motion in vector form reads


∂Φ 1 ∂Φ ∂Φ
~a = −∇Φ = ~eR + ~eφ + ~ez
∂R R ∂φ ∂z
Combining this with the above we see that

∂Φ v2
v̇R = − ∂R + Rφ
v v
v̇φ = − R1 ∂R
∂Φ
+ RR φ
v̇z = − ∂Φ∂z

which allows us to the write the CBE in cylindrical coordinates as


 2   
∂f ∂f vφ ∂f ∂f vφ ∂Φ ∂f 1 ∂Φ ∂f ∂Φ ∂f
+ vR + + vz + − − vR vφ + − =0
∂t ∂R R ∂φ ∂z R ∂R ∂vR R ∂φ ∂vφ ∂z ∂vz

The Jeans equations follow from multiplication with vR , vφ , and vz and integrat-
ing over velocity space. Note that the cylindrical symmetry requires that all
derivatives with respect to φ vanish. The remaining terms are:

120
Z Z
∂f ∂ ∂(ρhvR i)
vR d3~v = vR f d3~v =
∂t ∂t ∂t
Z Z
∂f 3 ∂ ∂(ρhvR2 i)
vR2 d ~v = vR2 f d3~v =
∂R ∂R ∂R
Z Z
∂f ∂ ∂(ρhvR vz i)
vR vz d3~v = vR vz f d3~v =
∂z ∂z ∂z
Z 2 Z 2 Z 
vR vφ ∂f 3 1 ∂(vR vφ f ) 3 ∂(vR vφ2 ) 3 hvφ2 i
d ~v = d ~v − f d ~v = −ρ
R ∂vR R ∂vR ∂vR R
Z Z Z 
∂Φ ∂f 3 ∂Φ ∂(vR f ) 3 ∂vR 3 ∂Φ
vR d ~v = d ~v − f d ~v = −ρ
∂R ∂vR ∂R ∂vR ∂vR ∂R
Z 2 Z Z 
vR vφ ∂f 3 1 ∂(vR2 vφ f ) 3 ∂(vR2 vφ ) 3 hv 2 i
d ~v = d ~v − f d ~v = −ρ R
R ∂vφ R ∂vφ ∂vφ R
Z Z Z 
∂Φ ∂f 3 ∂Φ ∂(vR f ) 3 ∂vz 3
vR d ~v = d ~v − f d ~v = 0
∂z ∂vz ∂z ∂vz ∂vR

Working out the similar terms for the other Jeans equations we finally obtain the
Jeans Equations in Cylindrical Coordinates:

 2 
∂(ρhvR i) ∂(ρhvR2 i) ∂(ρhvR vz i) hvR i − hvφ2 i ∂Φ
+ + +ρ + =0
∂t ∂R ∂z R ∂R

∂(ρhvφ i) ∂(ρhvR vφ i) ∂(ρhvφ vz i) hvR vφ i


+ + + 2ρ =0
∂t ∂R ∂z R
 
∂(ρhvz i) ∂(ρhvR vz i) ∂(ρhvz2 i) hvR vz i ∂Φ
+ + +ρ + =0
∂t ∂R ∂z R ∂z

These are 3 equations with 9 unknowns, which can only be solved if we make addi-
tional assumptions. In particular, one often makes the following assumptions:

1 System is static ⇒ the ∂t
-terms are zero and hvR i = hvz i = 0.
2 Velocity dispersion tensor is diagonal ⇒ hvi vj i = 0 (if i 6= j).
3 Meridional isotropy ⇒ hvR2 i = hvz2i = σR2 = σz2 ≡ σ 2 .

121
Under these assumptions we have 3 unknowns left: hvφ i, hvφ2 i, and σ 2 , and the Jeans
equations reduce to

 2 
∂(ρσ 2 ) σ − hvφ2 i ∂Φ
+ρ + =0
∂R R ∂R

∂(ρσ 2 ) ∂Φ
+ρ =0
∂z ∂z

Since we now only have two equations left, the system is still not closed. If from
the surface brightness we can estimate the mass density, ρ(R, z), and hence (using
the Poisson equation) the potential Φ(R, z), we can solve the second of these Jeans
equations for the meridional velocity dispersion:
Z∞
2 1 ∂Φ
σ (R, z) = ρ dz
ρ ∂z
z

and the first Jeans equation then gives the mean square azimuthal velocity
hvφ2 i = hvφ i2 + σφ2 :

∂Φ R ∂(ρσ 2 )
hvφ2 i(R, z) 2
= σ (R, z) + R +
∂R ρ ∂R

Thus, although hvφ2 i is uniquely specified by the Jeans equations, we don’t know how
it splits in the actual azimuthal streaming, hvφ i, and the azimuthal dispersion,
σφ2 . Additional assumptions are needed for this.

————————————————-

122
A similar analysis, but for a spherically symmetric system, using the spherical coordi-
nate system (r, θ, φ), gives the following Jeans equations in Spherical Symmetry

∂(ρhvr i) ∂(ρhvr2 i) ρ  2  ∂Φ
+ + 2hvr i − hvθ2 i − hvφ2 i + ρ =0
∂t ∂r r ∂r
∂(ρhvθ i) ∂(ρhvr vθ i) ρ   
+ + 3hvr vθ i + hvθ2 i − hvφ2 i cotθ = 0
∂t ∂r r
∂(ρhvφ i) ∂(ρhvr vφ i) ρ
+ + [3hvr vφ i + 2hvθ vφ icotθ] = 0
∂t ∂r r

If we now make the additional assumptions that the system is static and that also
the kinematic properties of the system are spherically symmetric, then there can
be no streaming motions and all mixed second-order moments vanish. Consequently,
the velocity dispersion tensor is diagonal with σθ2 = σφ2 . Under these assumptions
only one of the three Jeans equations remains:

∂(ρσr2 ) 2ρ  2  ∂Φ
+ σr − σθ2 + ρ =0
∂r r ∂r

Notice that this single equation still constains two unknown, σr2 (r) and σθ2 (r) (if we
assume that the density and potential are known), and can thus not be solved.
It is useful to define the anisotropy parameter

σθ2 (r) + σφ2 (r) σθ2 (r)


β(r) ≡ 1 − = 1 −
2σr2 (r) σr2 (r)

where the second equality only holds under the assumption that the kinematics are
spherically symmetric.

With β thus defined the (spherical) Jeans equation can be written as

1 ∂(ρhvr2 i) βhvr2 i dΦ
+2 =−
ρ ∂r r dr

If we now use that dΦ/dr = GM(r)/r, we can write the following expression for the

123
enclosed (dynamical) mass:
 
rhvr2 i d ln ρ d lnhvr2 i
M(r) = − + + 2β
G d ln r d ln r

Hence, if we know ρ(r), hvr2 i(r), and β(r), we can use the spherical Jeans equation
to infer the mass profile M(r).

Consider an external, spherical galaxy. Observationally, we can measure the pro-


jected surface brightness profile, Σ(R), which is related to the 3D luminosity
density ν(r) = ρ(r)/Υ(r)
Z∞
ν r dr
Σ(R) = 2 √
r 2 − R2
R

with Υ(r) the mass-to-light ratio. Similarly, the line-of-sight velocity disper-
sion, σp2 (R), which can be inferred from spectroscopy, is related to both hvr2 i(r) and
β(r) according to (see Fig. 19)

Z∞
ν r dr
Σ(R)σp2 (R) = 2 h(vr cos α − vθ sin α)2 i √
r 2 − R2
R
Z∞
 ν r dr
= 2 hvr2 i cos2 α + hvθ2 i sin2 α √
r 2 − R2
R
Z∞  
R2 ν hvr2 i r dr
= 2 1−β 2 √
r r 2 − R2
R

The 3D luminosity density is trivially obtained from the observed Σ(R) using the
Abel transform
Z∞
1 dΣ dR
ν(r) = − √
π dR R2 − r 2
r

In general, we have three unknowns: M(r) [or equivalently ρ(r) or Υ(r)], hvr2 i(r) and
β(r). With our two observables Σ(R) and σp2 (R), these can only be determined if we
make additional assumptions.

124
Figure 19: Geometry related to projection

EXAMPLE 1: Assume isotropy: β(r) = 0. In this case we can use the Abel
transform to obtain
Z∞
2 1 d(Σσp2 ) dR
ν(r)hvr i(r) = − √
π dR R2 − r 2
r

and the enclosed mass follows from the Jeans equation


 
rhvr2 i d ln ν d lnhvr2 i
M(r) = − +
G d ln r d ln r
Note that the first term uses the luminosity density ν(r) rather than the mass density
ρ(r). This is because σp2 is weighted by light rather than mass.
The mass-to-light ratio now follows from
M(r)
Υ(r) = Rr
4π 0 ν(r) r 2 dr
which can be used to constrain the mass of a potential dark matter halo or central
supermassive black hole (but always under assumption that the system is isotropic).

EXAMPLE 2: Assume a constant mass-to-light ratio: Υ(r) = Υ0 . In this case the


luminosity density ν(r) immediately yields the enclosed mass:
Zr
M(r) = 4πΥ0 ν(r) r 2 dr
0

125
We can now use the spherical Jeans Equation to write β(r) in terms of M(r),
ν(r) and hvr2 i(r). Substituting this in the equation for Σ(R)σp2 (R) yields a solution
for hvr2i(r), and thus for β(r). As long as β(r) ≤ 1 the model is said to be self-
consistent within the context of the Jeans equations.
Almost always, radically different models (based on radically different assumptions)
can be constructed, that are all consistent with the data and the Jeans equations.
This is often referred to as the mass-anisotropy degeneracy. Note, however, that
none of these models need to be physical: they can still have f < 0.

126
CHAPTER 16

Orbit Theory

The ‘building-blocks’ of collisionless systems, such as galaxies and dark matter halos,
are orbits. In this Chapter, we very briefly highlight a few aspects of orbit theory.
A more detailed account of orbit theory in the context of astrophysical systems can
be found in the excellent textbook ”Galactic Dynamics” by Binney & Tremaine.

Integrals of Motion: An integral of motion is a function I(~x, ~v) of the phase-space


coordinates that is constant along all orbits, i.e.,

I[~x(t1 ), ~v (t1 )] = I[~x(t2 ), ~v(t2 )]

for any t1 and t2 . The value of the integral of motion can be the same for different
orbits. Note that an integral of motion can not depend on time. Orbits can have
from zero to five integrals of motion. If the Hamiltonian does not depend on time,
then energy is always an integral of motion.

Integrals of motion come in two kinds:


• Isolating Integrals of Motion: these reduce the dimensionality of the particle’s
trajectory in 6-dimensional phase-space by one. Therefore, an orbit with n isolating
integrals of motion is restricted to a 6 − n dimensional manifold in 6-dimensional
phase-space. Energy is always an isolating integral of motion.
• Non-Isolating Integrals of Motion: these are integrals of motion that do not
reduce the dimensionality of the particle’s trajectory in phase-space. They are of
essentially no practical value for the dynamics of the system.

Orbits: If in a system with n degrees of freedom a particular orbit admits n in-


dependent isolating integrals of motion, the orbit is said to be regular, and its
corresponding trajectory Γ(t) is confined to a 2n − n = n dimensional manifold in
phase-space. Topologically this manifold is called an invariant torus (or torus for
short), and is uniquely specified by the n isolating integrals. A regular orbit has n
fundamental frequencies, ωi , with which it circulates or librates in its n-dimensional
manifold. If two or more of these frequencies are commensurable (i.e., lωi + mωj = 0
with l and m integers), then the orbit is a resonant orbit, and has a dimensionality

127
that is one lower than that of the non-resonant regular orbits (i.e., lωi + mωj is an
extra isolating integral of motion). Orbits with fewer than n isolating integrals of
motion are called irregular or stochastic.

Every spherical potential admits at least four isolating integrals of motion, namely
energy, E, and the three components of the angular momentum vector L. ~ Orbits in
a flattened, axisymmetric potential frequently (but not always) admit three isolating
integrals of motion: E, Lz (where the z-axis is the system’s symmetry axis), and a
non-classical third integral I3 (the integral is called non-classical since there is no
analytical expression of I3 as function of the phase-space variables).

Since an integral of motion, I(~x, ~v ) is constant along an orbit, we have that

dI ∂I dxi ∂I dvi ∂I
= + = ~v · ∇I − ∇Φ · =0
dt ∂xi dt ∂vi dt ∂~v
Compare this to the CBE for a steady-state (static) system:

∂f
~v · ∇f − ∇Φ · =0
∂~v
Thus the condition for I to be an integral of motion is identical with the condition
for I to be a steady-state solution of the CBE. This implies the following:

Jeans Theorem: Any steady-state solution of the CBE depends on the phase-
space coordinates only through integrals of motion. Any function of these integrals is
a steady-state solution of the CBE.

Strong Jeans Theorem: The DF of a steady-state system in which almost all


orbits are regular can be written as a function of the independent isolating integrals
of motion.

~
Hence, the DF of any steady-state spherical system can be expressed as f = f (E, L).
2
If the system is spherically symmetric in all its properties, then f = f (E, L ), i.e.,
the DF can only depend on the magnitude of the angular momentum vector, not on
its direction.

128
An even simpler case to consider is the one in which f = f (E): Since E = Φ(~r) +
1 2
[v + vθ2 + vφ2 ] we have that
2 r
Z  
2 1 2 1 2 2 2
hvr i = dvr dvθ dvφ vr f Φ + [vr + vθ + vφ ]
ρ 2
Z  
2 1 2 1 2 2 2
hvθ i = dvr dvθ dvφ vθ f Φ + [vr + vθ + vφ ]
ρ 2
Z  
2 1 2 1 2 2 2
hvφ i = dvr dvθ dvφ vφ f Φ + [vr + vθ + vφ ]
ρ 2

Since these equations differ only in the labelling of one of the variables of integration,
it is immediately evident that hvr2 i = hvθ2 i = hvφ2 i. Hence, assuming that f = f (E) is
identical to assuming that the system is isotropic (and thus β(r) = 0). And since
Z  
1 1 2 2 2
hvi i = dvr dvθ dvφ vi f Φ + [vr + vθ + vφ ]
ρ 2
it is also immediately evident that hvr i = hvθ i = hvφ i = 0. Thus, a system with
f = f (E) has no net sense of rotation.

The more general f (E, L2 ) models typically are anisotropic. Models with 0 < β ≤ 1
are radially anisotropic. In the extreme case of β = 1 all orbits are purely radial
and f = g(E) δ(L), with g(E) some function of energy. Tangentially anisotropic
models have β < 0, with β = −∞ corresponding to a model in which all orbits are
circular. In that case f = g(E) δ[L − Lmax (E)], where Lmax (E) is the maximum
angular momentum for energy E. Another special case is the one in which β(r) = β
is constant; such models have f = g(E) L−2β .

Next we consider axisymmetric systems. If we only consider systems for which


most orbits are regular, then the strong Jeans Theorem states that, in the most
general case, f = f (E, Lz , I3 ). For a static, axisymmetric system

hvR i = hvz i = 0 hvR vφ i = hvz vφ i = 0

but note that, in this general case, hvR vz i =


6 0; Hence, in general, in a three-integral
model with f = f (E, Lz , I3 ) the velocity ellipsoid is not aligned with (R, φ, z), and
the velocity dispersion tensor contains four unknowns: hvR2 i, hvφ2 i, hvz2 i, and hvR vz i.
In this case there are two non-trivial Jeans Equations:

129
 2 
∂(ρhvR2 i) ∂(ρhvR vz i) hvR i − hvφ2 i ∂Φ
+ +ρ + =0
∂R ∂z R ∂R
 
∂(ρhvR vz i) ∂(ρhvz2 i) hvR vz i ∂Φ
+ +ρ + =0
∂R ∂z R ∂z

which clearly doesn’t suffice to solve for the four unknowns (modelling three-integral
axisymmetric systems is best done using the Schwarzschild orbit superposition tech-
nique). To make progress with Jeans modeling, one has to make additional assump-
tions. A typical assumption is that the DF has the two-integral form f = f (E, Lz ).

In that case, hvR vz i = 0 [velocity ellipsoid now is aligned with (R, φ, z)] and hvR2 i =
hvz2 i (see Binney & Tremaine 2008), so that the Jeans equations reduce to

 2 
∂(ρhvR2 i) hvR i − hvφ2 i ∂Φ
+ρ + =0
∂R R ∂R

∂(ρhvz2 i) ∂Φ
+ρ =0
∂z ∂z

which is a closed set for the two unknowns hvR2 i (= hvz2 i) and hvφ2 i. Note, however,
that the Jeans equations provide no information about how hvφ2 i splits in streaming
and random motions. In practice one often writes that
 
hvφ i = k hvφ2 i − hvR2 i

with k a free parameter. When k = 1 the azimuthal dispersion is σφ2 ≡ hvφ2 i − hvφ i2 =
σR2 = σz2 everywhere. Such models are called oblate isotropic rotators.

130
CHAPTER 17

Collisions & Encounters of Collisionless Systems

Consider an encounter between two collisionless N-body systems (i.e., dark matter
halos or galaxies): a perturber P and a system S. Let q denote a particle of S and
let b be the impact parameter, v∞ the initial speed of the encounter, and R0 the
distance of closest approach (see Fig. 20).

Typically what happens in an encounter is that orbital energy (of P wrt S) is


converted into random motion energy of the constituent particles of P and S
(i.e., q gains kinetic energy wrt S).
R
The velocity impulse ∆~vq = ~g (t) dt of q due to the gravitational field ~g (t) from
P decreases with increasing v∞ (simply because ∆t will be shorter). Consequently,
when v∞ increases, less and less orbital energy is transferred to random motion, and
there is a critical velocity, vcrit , such that
v∞ > vcrit ⇒ S and P escape from each other
v∞ < vcrit ⇒ S and P merge together

There are only two cases in which we can calculate the outcome of the encounter
analytically:

• high speed encounter (v∞ ≫ vcrit ). In this case the encounter is said to
be impulsive and one can use the impulsive approximation to compute its
outcome.

• large mass ratio (MP ≪ MS ). In this case one can use the treatment of
dynamical friction to describe how P loses orbital energy and angular mo-
mentum to S.

In all other cases, one basically has to resort to numerical simulations to study the
outcome of the encounter. In what follows we present treatments of first the impulse
approximation and then dynamical friction.

131
Figure 20: Schematic illustration of an encounter with impact parameter b between
a perturber P and a subject S.

Impulse Approximation: In the limit where the encounter velocity v∞ is much


larger than the internal velocity dispersion of S, the change in the internal energy of
S can be approximated analytically. The reason is that, in a high-speed encounter,
the time-scale over which the tidal forces from P act on q is much shorter than the
dynamical time of S (or q). Hence, we may consider q to be stationary (fixed wrt
S) during the encounter. Consequently, q only experiences a change in its kinetic
energy, while its potential energy remains unchanged:
1 1 1
∆Eq = (~v + ∆~v )2 − ~v 2 = ~v · ∆~v + |∆~v |2
2 2 2
P
We are interested in ∆ES = q ∆Eq , where the summation is over all its constituent
particles:
Z Z
3 1
∆ES = ∆E(~r) ρ(r) d ~r ≃ |∆~v |2 ρ(r)d3~r
2
where we have used that, because of symmetry, the integral
Z
~v · ∆~v ρ(r) d3~r ≃ 0

In the large v∞ limit, we have that the distance of closest approach R0 → b, and the
velocity of P wrt S is vP (t) ≃ v∞~ey ≡ vP~ey . Consequently, we have that

~
R(t) = (b, vP t, 0)

132
Let ~r be the position vector of q wrt S and adopt the distant encounter approx-
imation, which means that b ≫ max[RS , RP ], where RS and RP are the sizes of S
and P , respectively. This means that we may treat P as a point mass MP , so that
GMP
ΦP (~r) = −
~
|~r − R|
~ we we have that
Using geometry, and defining φ as the angle between ~r and R,

~ 2 = (R − r cos φ)2 + (r sin φ)2


|~r − R|
so that
p
~ =
|~r − R| R2 − 2rR cos φ + r 2
Next we use the series expansion
1 1 13 2 135 3
√ =1− x+ x − x + ....
1+x 2 24 246
to write
"    2 #
1 1 1 r r2 3 r r2
= 1− −2 cos φ + 2 + −2 cos φ + 2 + ...
~
|~r − R| R 2 R R 8 R R
Substitution in the expression for the potential of P yields
 
GMP GMP r GMP r 2 3 2 1
ΦP (~r) = − − cos φ − cos φ − + O[(r/R)3 ]
R R2 R3 2 2

• The first term on rhs is a constant, not yielding any force (i.e., ∇r ΦP = 0).
• The second term on the rhs describes how the center of mass of S changes its
velocity due to the encounter with P .
• The third term on the rhs corresponds to the tidal force per unit mass and is
the term of interest to us.

It is useful to work in a rotating coordinate frame (x′ , y ′, z ′ ) centered on S and with


the x′ -axis pointing towards the instantaneous location of P , i.e., x′ points along
~
R(t). Hence, we have that x′ = r ′ cos φ, where r ′ 2 = x′ 2 + y ′ 2 + z ′ 2 . In this new
coordinate frame, we can express the third term of ΦP (~r) as

133
 
GMP 3 ′ 2 2 1 ′2
Φ3 (~r) = − 3 r cos φ − r
R 2 2
 
GMP ′2 1 ′2 1 ′2
= − 3 x − y − z
R 2 2

Hence, the tidal force is given by


GMP
F~tid

(~r) ≡ −∇ΦP = (2x′ , −y ′ , −z ′ )
R3
We can relate the components of F~tid ′
to those of the corresponding tidal force, F~tid
in the (x, y, z)-coordinate system using

x′ = x cos θ − y sin θ Fx = Fx′ cos θ + Fy′ sin θ


y ′ = x sin θ + y cos θ Fy = −Fx′ sin θ + Fy′ cos θ
z′ = z Fz = Fz ′
where θ is the angle between the x and x′ axes, with cos θ = b/R and sin θ = vP t/R.
After some algebra one finds that

GMP  2

Fx = x(2 − 3 sin θ) − 3y sin θ cos θ
R3
GMP  2

Fy = y(2 − 3 cos θ) − 3x sin θ cos θ
R3
GMP
Fx = − 3 z
R
Using these, we have that
Z Z Z π/2
dvx dt
∆vx = dt = Fx dt = Fx dθ
dt −π/2 dθ
with similar expressions for ∆vy and ∆vz . Using that θ = tan−1 (vP t/b) one has that
dt/dθ = b/(vP cos2 θ). Substituting the above expressions for the tidal force, and
using that R = b/ cos θ, one finds, after some algebra, that
2GMP
∆~v = (∆vx , ∆vy , ∆vz ) = (x, 0, −z)
vP b2
Substitution in the expression for ∆ES yields

134
Z
1 2 G2 MP2
∆ES = |∆~v |2 ρ(r) d3~r = MS hx2 + z 2 i
2 vP2 b4
Under the assumption that S is spherically symmetric we have that hx2 + z 2 i =
2
3
hx2 + y 2 + z 2 i = 23 hr 2 i and we obtain the final expression for the energy increase of
S as a consequence of the impulsive encounter with P :
 2
4 MP hr 2 i
∆ES = G2 MS
3 vP b4

This derivation, which is originally due to Spitzer (1958), is surprisingly accurate for
encounters with b > 5max[RP , RS ], even for relatively slow encounters with v∞ ∼ σS .
For smaller impact parameters one has to make a correction (see Galaxy Formation
and Evolution by Mo, van den Bosch & White 2010 for details).

The impulse approximation shows that high-speed encounters can pump energy
into the systems involved. This energy is tapped from the orbital energy of the two
systems wrt each other. Note that ∆ES ∝ b−4 , so that close encounters are far more
important than distant encounters.

Let Eb ∝ GMS /RS be the binding energy of S. Then, it is tempting to postulate


that if ∆ES > Eb the impulsive encounter will lead to the tidal disruption of S.
However, this is not at all guaranteed. What is important for disruption is how that
energy ∆ES is distributed over the consistituent particles of S. Since ∆E ∝ r 2 ,
particles in the outskirts of S typically gain much more energy than those in the
center. However, particles in the outskirts are least bound, and thus require the
least amount of energy to become unbound. Particles that are strongly bound (with
small r) typically gain very little energy. As a consequence, a significant fraction
of the mass can remain bound even if ∆ES ≫ Eb (see van den Bosch et al., 2019,
MNRAS, 474, 3043 for details).

After the encounter, S has gained kinetic energy (in the amount of ∆ES ), but
its potential energy has remained unchanged (recall, this is the assumption that
underlies the impulse approximation). As a consequence, after the encounter S will
no longer be in virial equilibrium; S will have to readjust itself to re-establish
virial equilibrium.

135
Let K0 and E0 be the initial (pre-encounter) kinetic and total energy of S. The
virial theorem ensures that E0 = −K0 . The encounter causes an increase of
(kinetic) energy, so that K0 → K0 + ∆ES and E0 → E0 + ∆ES . After S has re-
established virial equilibrium, we have that K1 = −E1 = −(E0 + ∆ES ) = K0 − ∆ES .
Thus, we see that virialization after the encounter changes the kinetic energy of
S from K0 + ∆ES to K0 − ∆ES ! The gravitational energy after the encounter
is W1 = 2E1 = 2E0 + 2∆ES = W0 + 2∆ES , which is less negative than before
the encounter. Using the definition of the gravitational radius (see Chapter 18),
rg = GMS2 /|W |, from which it is clear that the (gravitational) radius of S increases
due to the impulsive encounter. Note that here we have ignored the complication
coming from the fact that the injection of energy ∆ES may result in unbinding some
of the mass of S.

Dynamical Friction: Consider the motion of a subject mass MS through a medium


of individual particles of mass m ≪ MS . The subject mass MS experiences a ”drag
force”, called dynamical friction, which transfers orbital energy and angular momen-
tum from MS to the sea of particles of mass m.

There are three different ”views” of dynamical friction:

1. Dynamical friction arises from two-body encounters between the subject


mass and the particles of mass m, which drives the system twowards equipar-
tition. i.e., towards 21 MS vS2 = 21 mhvm
2
i. Since MS ≫ m, the system thus
evolves towards vS ≪ vm (i.e., MS slowes down).

2. Due to gravitational focussing the subject mass MS creates an overdensity


of particles behind its path (the ”wake”). The gravitational back-reaction of
this wake on MS is what gives rise to dynamical friction and causes the subject
mass to slow down.

3. The subject mass MS causes a perturbation δΦ in the potential of the collection


of particles of mass m. The gravitational interaction between the response
density (the density distribution that corresponds to δΦ according to the
Poisson equation) and the subject mass is what gives rise to dynamical friction
(see Fig. 21).

Although these views are similar, there are some subtle differences. For example,
according to the first two descriptions dynamical friction is a local effect. The

136
third description, on the other hand, treats dynamical friction more as a global
effect. As we will see, there are circumstances under which these views make different
predictions, and if that is the case, the third and latter view presents itself as the
better one.

Chandrasekhar derived an expression for the dynamical friction force which, although
it is based on a number of questionable assumptions, yields results in reasonable
agreement with simulations. This so-called Chandrasekhar dynamical friction
force is given by

d~vS 4π G 2 MS2 ~vS


F~df = MS =− 2
ln Λ ρ(< vS )
dt vS vS

Here ρ(< vS ) is the density of particles of mass m that have a speed vm < vS , and ln Λ
is called the Coulomb logarithm. It’s value is uncertain (typically 3 ∼ < ln Λ < 30).

One often approximates it as ln Λ ∼ ln(Mh /MS ), where Mh is the total mass of
the system of particles of mass m, but this should only be considered a very rough
estimate at best. The uncertainties for the Coulomb logarithm derive from the
oversimplified assumptions made by Chandrasekhar, which include that the medium
through which the subject mass is moving is infinite, uniform and with an isotropic
velocity distribution f (vm ) for the sea of particles.

Similar to frictional drag in fluid mechanics, F~df is always pointing in the direction
opposite of vS .

Contrary to frictional drag in fluid mechanics, which always increases in strength


when vS increases, dynamical friction has a more complicated behavior: In the low-vS
limit, Fdf ∝ vS (similar to hydrodynamical drag). However, in the high-vS limit one
has that Fdf ∝ vS−2 (which arises from the fact that the factor ρ(< vS ) saturates).

Note that F~df is independent of the mass m of the constituent particles, and pro-
portional to MS2 . The latter arises, within the second or third view depicted above,
from the fact that the wake or response density has a mass that is proportional to
MS , and the gravitational force between the subject mass and the wake/response
density therefore scales as MS2 .

One shortcoming of Chandrasekhar’s dynamical friction description is that it treats

137
Figure 21: Examples of the response density in a host system due to a perturber
orbiting inside it. The back-reaction of this response density on the perturber causes
the latter to experience dynamical friction. The position of the perturber is indicated
by an asterisk. [Source: Weinberg, 1989, MNRAS, 239, 549]

dynamical friction as a purely local phenomenon; it is treated as the cumulative


effect of many uncorrelated two-body encounters between the subject masss and
the individual field particles. That this local treatment is incomplete is evident from
the fact that an object A orbiting outside of an N-body system B still experiences
dynamical friction. This can be understood with the picture sketched under view 3
above, but not in a view that treats dynamical friction as a local phenomenon.

Orbital decay: As a consequence of dynamical friction, a subject mass MS orbiting


inside (or just outside) of a larger N-body system of mass Mh > MS , will transfer
its orbital energy and angular momentum to the constituent particles of the ‘host’
mass. As a consequence it experiences orbital decay.

Let us assume that the host mass is a singular isothermal sphere with density
and potential given by

Vc2
ρ(r) = 2
Φ(r) = Vc2 ln r
4πGr
2
where Vc = GMh /rh with rh the radius of the host mass. If we further assume
that this host mass has, at each point, an isotropic and Maxwellian velocity
distrubution, then

138
 2

ρ(r) vm
f (vm ) = exp − 2
(2πσ 2 )3/2 2σ

with σ = Vc / 2.

NOTE: the assumption of a singular isothermal sphere with an isotropic, Maxwellian


velocity distribution is unrealistic, but it serves the purpose of the order-of-magnitude
estimate for the orbital decay rate presented below.

Now consider a subject of mass MS moving on a circular orbit (vS = Vc ) through


this host system of mass Mh . The Chandrasekhar dynamical friction that this
subject mass experiences is
 
4π ln Λ G 2 MS2 ρ(r) 2 −1 GMS2
Fdf = − erf(1) − √ e ≃ −0.428 ln Λ
Vc2 π r2
The subject mass has specific angular momentum L = rvS , which it loses due to
dynamical friction at a rate
dL dvS Fdf GMS
=r =r ≃ −0.428 ln Λ
dt dt MS r
Due to this angular momentum loss, the subject mass moves to a smaller radius,
while it continues to move on a circular orbit with vS = Vc . Hence, the rate at which
the orbital radius changes obeys
dr dL GMS
Vc = = −0.428 ln Λ
dt dt r
Solving this differential equation subject to the initial condition that r(0) = ri , one
finds that the subject mass MS reaches the center of the host after a time
 2
1.17 ri2 Vc 1.17 ri Mh rh
tdf = =
ln Λ G MS ln Λ rh MS Vc
In the case where the host system is a virialized dark matter halo we have that
rh 1
≃ = 0.1tH
Vc 10H(z)
where tH is called the Hubble time, and is approximately equal to the age of
the Universe corresponding to redshift z (the above relation derives from the fact

139
that virialized dark matter halos all have the same average density). Using that
ln Λ ∼ ln(Mh /MS ) and assuming that the subject mass starts out from an initial
radius ri = rh , we obtain a dynamical friction time

Mh /MS
tdf = 0.12 tH
ln(Mh /MS )
Hence, the time tdf on which dynamical friction brings an object of mass MS moving
in a host of mass Mh from an initial radius of ri = rh to r = 0 is shorter than the
Hubble time as long as MS ∼ > M /30. Hence, dynamical friction is only effective for
h
fairly massive objects, relative to the mass of the host. In fact, if you take into
account that the subject mass experiences mass stripping as well (due to the tidal
interactions with the host), the dynamical friction time increases by a factor 2 to 3,
and tdf < tH actually requires that MS ∼> M /10.
h

For a more detailed treatment of collisions and encounters of collisionless systems,


see Chapter 12 of ”Galaxy Formation and Evolution” by Mo, van den Bosch &
White.

140
Part III: Radiative Processes

With the exception of meteorites, neutrinos, gravitational waves, and cosmic rays, all
information about the Universe reaches us in the form of radiation. Understanding
how radiation is produced, and how it interacts with matter on its way from the
source to our telescopes is therefore of crucial importance for astrophysics. The
following chapters give an elementary introduction into these topics.

Radiative processes is a rich topic, and one could easily devote an entire course to it.
The following chapters therefore only scratch the surface of this rich topic. Readers
who want to get more indepth information are referred to the following excellent
textbooks
- Radiative Processes in Astrophysics by G. Rybicki & A. Lightman
- Astrophysics: Decoding the Cosmos by J. Irwin
- The Physics of Astrophysics I. Radiation by F. Shu
- Theoretical Astrophysics by M. Bartelmann

141
CHAPTER 18

Radiation Essentials

Spectral Energy Distribution: the radiation from a source may be characterized


by its spectral energy distribution (SED), Lν dν, or, equivalently, Lλ dλ. Some texts
refer to the SEDs as the spectral luminosity or the spectral power. The SED is the
total energy emitted by photons in the frequency interval [ν, ν + dν], and is related
to the total luminosity, L ≡ dE/dt, according to
Z Z
L = Lν dν = Lλ dλ

Note that [Lν ] = erg s−1 Hz−1 , while [L] = erg s−1 .

Flux: The flux, f , of a source is the radiation energy per unit time passing through
a unit area
dL = f dA [f ] = erg s−1 cm−2
where A is the area. Similarly, we can also define the spectral flux density (or
simply ‘flux density’), as the flux per unit spectral bandwidth:
dLν = fν dA [fν ] = erg s−1 cm−2 Hz−1
In radio astronomy, one typically expresses fν in Jansky, where 1Jy = 10−23 erg s−1 cm−2 Hz−1 .
As with the SEDs, one may also express spectral flux densities as fλ . Using that
λ = c/ν, and using that fν dν = fλ dλ one has that
λ2 ν2
fν = fλ , fλ = fν
c c

Luminosity and flux are related according to


L = 4 π r2 f
where r is the distance from the source.
Intensity: The intensity, I, also called surface brightness is the flux emitted in,
or observed from, a solid angle dΩ. The intensity is related to the flux via
df = I cos θ dΩ

142
Figure 22: Diagrams showing intensity and its dependence on direction and solid
angle. Fig. (a) depicts the ‘observational view’, where dA represents an element of
a detector. The arrows show incoming rays from the center of the source. Fig. (b)
depicts the ‘emission view’, where dA represents the surface of a star. At each point
on the surface, photons leave in all directions away from the surface.

where θ is the angle between the normal of the surface area through which the flux
is measured and the direction of the solid angle. The unit of intensity is [I] =
erg s−1 cm−2 sr−1 . Here ‘sr’ is a steradian, which is the unit of solid angle measure
(there are 4π steradians in a complete sphere). As with the flux and luminosity,
one can also define a specific intensity, Iν , which is the intensity per unit spectral
bandwidth ([Iν ] = erg s−1 cm−2 Hz−1 sr−1 ).

The flux emerging from the surface of a star with luminosity L and radius R∗ is
Z Z 2π Z π/2
L
F ≡ = I cos θ dΩ = dφ dθ I cos θ sin θ = π I
4πR∗2 half sphere 0 0

where we have used that dΩ = sin θ dθ dφ, and the fact that the integration over the
solid angle Ω is only to be performed over half a sphere. Note that an observer can
only measure the surface brightness of resolved objects; if unresolved, the observer
can only measure the objects flux.

Consider a resolved object (i.e., a galaxy), whose surface brightness distribution on


the sky is given by I(Ω). If the objects extents a solid angle ΩS on the sky, its flux

143
is given by Z Z
f= I(Ω) cos θ dΩ ≃ I(Ω) dΩ ≡ hIi ΩS
ΩS

where we have assumed that ΩS is small, so that variations of cos θ across the object
can be neglected. Since both f ∝ r −2 and ΩS ∝ r −2 , where r is the object’s distance,
we see that the average surface brightness hIi is independent of distance.

Energy density: the energy density, u, is a measure of the radiative energy per unit
volume (i.e., [u] = erg cm−3 ). If the radiation intensity as seen from some specific
location in space is given by I(Ω), then the energy density at that location is
Z
1 4π
u= I dΩ ≡ J
c c
where Z
1
J≡ I dΩ

is the mean intensity (i.e., average over 4π sterradian). If the radiation is isotropic
(i.e., the center of a star, or, to good approximation, a random location in the early
Universe), then J = I. If the radiation intensity
P is due to the summed intensity from
1
a number of individual sources, then u = c i fi , where fi is the flux due to source
i.

Recall from Chapter 6 that the number density of photons emerging from a Black
Body of temperature T is given by

8π ν 2 dν
nγ (ν, T ) dν = 3 hν/k
c e BT − 1

Hence, we have that

8π h ν 3 dν
u(ν, T ) dν = nγ (ν, T ) hν dν =
c3 ehν/kB T − 1
Using that u(ν, T ) = (4π/c)Jν (T ) we have that the mean specific intensity from
a black body [for which one typically uses the symbol Bν (T )] is given by

2 h ν3 dν
Bν (T ) dν = 2 hν/k
c e BT − 1

144
Figure 23: Various Planck curves for different temperatures, illustrating Wien’s dis-
placement law. Note how the Planck curve for a black body with the temperature of
the Sun peaks at the visible wavelengths, where the sensitivity of our eyes is maximal

which is called the Planck curve (or ‘formula’). Integrating over frequency yields
the total, mean intensity emitted from the surface of a Black Body
Z ∞
σSB 4
J = J(T ) = Bν (T ) dν = T
0 π
where σSB is the Stefan-Boltzmann constant. This implies an energy density
4π 4σSB 4
u = u(T ) = J= T ≡ ar T 4
c c
where ar ≃ 7.6 × 10−15 erg cm−3 K−4 is called the radiation constant (see also
Chapter 6).

Wien’s Displacement Law: When the temperature of a Black Body emitter in-
creases, the overall radiated energy increases and the peak of the radiation curve
moves to shorter wavelengths. It is straightforward to show that the product of the
temperature and the wavelength at which the Planck curve peaks is a constant, given
by
λmax T = 0.29

145
where T is the absolute temperature, expressed in degrees Kelvin, and λmax is ex-
pressed in cm. This relation is called Wien’s Displacement Law.

Stefan-Boltzmann Law: The flux emitted by a Black Body is

FBB = π I(T ) = σSB T 4

which is known as the Stefan-Boltzmann law. This law is used to define the
effective temperature of an emitter.

Effective Temperature: The temperature an emitter of flux F would have if it


where a Black Body; using the Stefan-Boltzmann law we have that Teff = (F/σSB )1/4 .
We can also use the effective temperature to express the emitter’s luminosity;

L = 4 π R2 σSB Teff
4

where R is the emitter’s radius. The effective temperature is sometimes also called
the radiation temperature, as a measure for the temperature associated with the
radiation field.

Brightness Temperature: the brightness temperature, TB (ν), of a source at fre-


quency ν is defined as the temperature which, when put into the Planck formula,
yields the specific intensity actually measured at that frequency. Hence, for a Black
Body TB (ν) is simply equal to the temperature of the Black Body. If TB (ν) depends
on frequency, then the emitter is not a Black Body. The brightness temperature is
a frequency-dependent version of the effective, or radiation, temperature.

Wavebands: Astronomers typically measure an object’s flux through some filter


(waveband). The measured flux in ‘band’ X is, fX , is related to the spectral flux
density, fλ , of the object according to
Z
fX = fλ FX (λ) R(λ) T (λ) dλ

Here FX (λ) describes the transmission of the filter that defines waveband X, R(λ)
is the transmission efficiency of the telescope + instrument, and T (λ) describes the
transmission of the atmosphere. The combined effect of FX , R, and T is typically
‘calibrated’ using standard stars with known fλ .

146
Magnitudes: For historical reasons, the flux of an astronomical object in waveband
X is usually quoted in terms of apparent magnitude:
 
fX
mX = −2.5 log
fX,0

where the flux zero-point fX,0 has traditionally been taken as the flux in the X
band of the bright star Vega. In recent years it has become more common to use
‘AB-magnitudes’, for which
Z
−20 −1 −2 −1
fX,0 = 3.6308 × 10 erg s cm Hz FX (c/ν) dν

Similarly, the luminosities of objects (in waveband X) are often quoted as an abso-
lute magnitude:
MX = −2.5 log(LX ) + CX
where CX is a zero point. It is usually convenient to write LX in units of the solar
luminosity in the same band, L⊙X , so that
 
LX
MX = −2.5 log + M⊙X ,
L⊙X

where M⊙X is the absolute magnitude of the Sun in the waveband in consideration.
Using the relation between luminosity and flux we have that

mX − MX = 5 log(r/r0 )

where r0 is a fiducial distance at which mX and MX are defined to have the same
value. Conventionally, r0 is chosen to be 10 pc.

Distance modulus: the distance modulus of an object is defined as mX − MX .

147
CHAPTER 19

Thermal Equilibrium & Saha Equation

Most of the baryonic matter in the Universe is in a gaseous state, made up of ∼ 75%
Hydrogen (H), ∼ 25% Helium (He) and only small amounts of other elements (called
‘metals’). Gases can be neutral, ionized, or partially ionized. The degree of ionization
of any given element is specified by a Roman numeral after the element name. For
example HI and HII refer to neutral and ionized hydrogen, respectively, while CI is
neutral carbon, CII is singly-ionized carbon, and CIV is triply-ionized carbon. A gas
that is highly (largely) ionized is called a plasma.

Thermodynamic Equilibrium: a system is said to be in thermodynamic equilib-


rium (TE) when it is in
- thermal equilibrium
- mechanical equilibrium
- radiative equilibrium
- chemical equilibrium
- statistical equilibrium
Equilibrium means a state of balance. In a state of thermodynamic equilibrium,
there are no net flows of matter or of energy, no phase changes, and no unbalanced
potentials (or driving forces), within the system. A system that is in thermodynamic
equilibrium experiences no changes when it is isolated from its surroundings.

In a system in TE the energy in the radiation field is in equilibrium with the kinetic
energy of the particles. If the system is isolated (to both matter and radiation), and
in mechanical equilibrium, then over time TE will be established. For a gas in TE,
the radiation temperature, TR , is equal to the kinetic temperature, T , is equal
to the excitation temperature, Tex (see below for definitions). Since no photons
are allowed to escape from a system in TE (this would correspond to energy loss,
and thus violate TE), the photons that are produced in the gas (represented by TR )
therefore must be tightly coupled to the random motions of the particles (represented
by T ). This coupling implies that, for bound states, the excitation and de-excitation
of the atoms and ions must be dominated by collisions. In other words, the collision
timescales must be shorter than the timescales associated with photon interactions

148
or spontaneous de-excitations. If that were not the case, T could not remain
equal to TR .

Local Thermodynamic Equilibrium (LTE): True TE is rare (in almost all cases
energy does escape the system in the form of radiation, i.e., the system cools), and
often temperature gradients are present. A good, albeit imperfect, example of TE
is the Universe as a whole prior to decoupling. Although true TE is rare, in many
systems (stars, gaseous spheres, ISM), we may apply local TE (LTE), which implies
that the gas is in TE, but only locally. In a system in LTE, there will typically be
gradients in temperature, density, pressure, etc, but they are sufficiently small over
the mean-free path of a gas particle. For stellar interiors, the fact that radiation is
‘locally trapped’ explains why it takes so long for photons to diffuse from the center
(where they are created in nuclear reactions) to the surface (where they are emitted
into space). In the case of the Sun, this timescale is of the order of 200,000 years.

Thermal Equilibrium: For a system to be in thermal equilibrium with itself, it


must not contain temperature gradients. For two systems to be in thermal equilib-
rium with each other requires an actual or implied thermal connection between them,
through a path that is permeable only to heat, and that no net energy is transferred
through that path.

For a gas in thermal equilibrium at some uniform temperature, T , and uniform


number density, n, the number density of particles with speeds between v and v + dv
is given by the Maxwell-Boltzmann (aka ‘Maxwellian’) velocity distribution:
 3/2  
m m v2
n(v) dv = n exp − 4πv 2 dv
2π kB T 2 kB T

where m is the mass of a gas particle. The temperature T is called the kinetic
temperature, and is related to the mean-square particle speed, hv 2i, according to
1 3
m hv 2i = kB T
2 2
p
The most probable speed of the Maxwell-Boltzmann
p distribution is vmp = 2kB T /m,
while the mean speed is vmean = 8kB T /m.

149
Setting up a Maxwellian velocity distribution requires many elastic collisions. In
the limit where most collisions are elastic, the system will typically very quickly equi-
librate to thermal equilibrium. However, depending on the temperature of the gas,
collisions can also be inelastic: examples of the latter are collisional excitations,
in which part of the kinetic energy of the particle is used to excite its target to an
excited state (i.e., the kinetic energy is now temporarily stored as a potential energy).
If the de-excitation is collisional, the energy is given back to the kinetic energy of
the gas (i.e., no photon is emitted). However, if the de-excitation is spontaneous
or via stimulated emission, a photon is emitted. If the gas is optically thin to the
emitted photon, the energy will escape the gas. The net outcome of the collisional
excitation is then one of dissipation, i.e., cooling (kinetic energy of the gas is being
radiated away). In the optically thick limit, the photon will be absorbed by another
atom (or ion). If the system is in (L)TE, the radiation temperature of these
photons being emitted and absorbed is equal to the kinetic temperature of the
gas. Although in LTE a subset of the collisions are inelastic, this subset is typically
small, and one may still use the Maxwell-Boltzmann distribution to characterize the
velocities of the gas particles.

Mechanical Equilibrium: A system is said to be in mechanical equilibrium if


(i) the vector sum of all external and internal forces is zero and (ii) the sum of the
moments of all external and internal forces about any line is zero. The internal forces
arise from pressure gradients. If a system is out of mechanical equilibrium, it will
try to re-establish mechanical equilibrium, the characteristic time-scale of which is
the sound-crossing time.

Radiative Equilibrium: A system is said to be in radiative equilibrium if any two


randomly selected subsystems of it exchange by radiation equal amounts of heat with
each other.

Chemical Equilibrium: In a chemical reaction, chemical equilibrium is the state


in which both reactants and products are present at concentrations which have no
further tendency to change with time. Usually, this state results when the forward
reaction proceeds at the same rate as the reverse reaction. The reaction rates of the
forward and reverse reactions are generally not zero but, being equal, there are no
net changes in the concentrations of the reactant and product.

150
Statistical Equilibrium: A system is said to be in statistical equilibrium if the
level populations of its constituent atoms and ions do not change with time (i.e., if
the transition rate into any given level equals the rate out).

For a system in statistical equilibrium, the population of states is described by the


Boltzmann law for level populations:
 
Nj gj h νij
= exp −
Ni gi kB T

Here Ni is the number of atoms in which electrons are in energy level i (here i reflects
the principal quantum number, with i = 1 refering to the ground state), νij is the
frequency corresponding to the energy difference ∆Eij = h νij of the energy levels, h
is the Planck constant, and gi is the statistical weight of population i.

Statistical weight: the statistical weight (aka ‘degeneracy’) indicates the number
of states at a given principal quantum number, n. In quantum mechanics, a total
of four quantum numbers is needed to describe the state of an electron: the prin-
ciple quantum number, n, the orbital angular momentum quantum number, l, the
magnetic quantum number ml , and the electron spin quantum number, ms . For
hydrogen l can take on the values 0, 1, ..., n − 1, the quantum number ml can take on
the values −l, −l + 1, −1, 0, 1, ..., l, and ms can take on the values +1/2 (‘up’) and
−1/2 (‘down’). Hence, gn = 2 n2 .

Excitation Temperature: The above Boltzmann law defines the excitation tem-
perature, Tex , as the temperature which, when put into the Boltzmann law for level
populations, results in the observed ratio of Nj /Ni . For a gas in (local) TE, all levels
in an atom can be described by the same Tex , which is also equal to the kinetic
temperature, T , and the radiation temperature, TR . Under non-LTE conditions,
each pair of energy levels can have a different Tex .

Radiation Temperature: the radiation temperature, or TR , is a temperature that


specifies the energy density in the radiation field, uγ , according to uγ = ar TR4 . In TE
TR = T , while for a Black Body, TR is the temperature that appears in the Planck
function that describes the specific intensity of the radiation field (see Appendix H
for more details).

151
Figure 24: Illustration of the energy levels of the hydrogen atom and some of the
most important transitions.

Partition Function: Using the above Boltzmann law, it is straightforward to


show that  
Nn gn ∆En
= exp −
N U kB T
P
where N = Ni is the total number of atoms (of all energy levels) and ∆En is the
energy difference between state n and the ground-state, and
X
U= gi e−∆Ei /kB T

is called the partition function. Note that at low temperatures, below that needed
to put a significant fraction of atoms in the first excited state, the partition function
becomes equal to the statistical weight of the ground state: U = g1 exp −0/kB T =
g1 = 2. After all, the exponential factors for the excited states are all extremely
small. (Note: the summation is over all principal quantum numbers up to some
maximum nmax , which is required to prevent divergence; see Irwin 2007 or Rybicki
& Lightmann 1979 for details).

Saha equation: the Saha equation expresses the ratios of atoms/ions in different
ionization states. In particular, the number of atoms/ions in the (K + 1)th ionization

152
Figure 25: The ionization fraction of hydrogen as a function of temperature, computed
using the Saha equation.

state versus those in the K th ionization state is given by


 3/2  
NK+1 2 UK+1 2 π me kB T χK
= exp −
NK ne Uk h2 kB T

where Un is the partition function of the nth ionization state, ne is the electron
number density, and χK is the energy required to remove an electron from the ground
state of the K th ionization state. Note that unlike the Boltzmann equation, the
Saha equation for the ionization fractions has a dependence on electron density. This
reflects that if there is a higher density of free electrons, there is a greater probability
that an electron will recombine with the ion, lowering the ionization state of the gas.

Hydrogen Since Hydrogen is the most common element in the Universe, it is impor-
tant to have some understanding of its structure. Fig. 18 shows the energy levels of
a hydrogen atom and some if its most important transitions. The Balmer transitions
have wavelengths that fall in the optical, and are therefore well known to (optical)
astronomers. The Lyman lines typically fall in the UV, and can only be observed
from space (or for high-redshift objects, for which the rest-frame UV is redshifted
into the optical).

Using the Saha equation, we can compute the ionization fraction for hydrogen
as a function of temperature and electron density. The ionization fraction is x =

153
NHII /[NHI + NHII ], which can be computed using the Saha equation:
3/2
 
NHII 15 T 1.58 × 105
= 2.41 × 10 exp −
NHI ne T

where we have used that χHI = 13.6eV, UHI = 2 and UHII = 1 (i.e., the ionized
hydrogen atom is just a free proton and only exists in a single state). Figure . 24
shows the ionization fraction x as function of temperature. Note that the transition
from almost neutral to almost completely ionized is extremely rapid! We can also
compute, using the Boltzmann law, the ratio of hydrogen atoms in the first excited
state compared to those in the ground state. The latter shows that the temperature
must exceed 30,000K in order for there to be an appreciable (10 percent) number of
hydrogen atoms in the first excited state. However, at such high temperatures, one
typically has that most of the hydrogen will be ionized (unless the electron density
is unrealistically high). This somewhat unintuitive result arises from the fact that
there are many more possible states available for a free electron than for a bound
electron in the first excited state. In conclusion, neutral hydrogen in LTE will have
virtually all its atoms in the ground state.

If densities are sufficiently low, which is typically the case in the ISM, the collisional
excitation rate (which scales with n2e ) is lower than the spontaneous de-excitation
rate. If that is the case, the hydrogen gas is no longer in LTE, and once again,
virtually all neutral hydrogen will find itself in the ground state. Neutral hydrogen
in the ISM is in the ground-state and is typically NOT in LTE. One important con-
sequence of the fact that neutral hydrogen is basically always observed in the ground
state, is that observations of hydrogen emission lines (Lyman, Balmer, Paschen, etc)
indicates that the hydrogen must be ionized; the lines arise from recombinations,
and are therefore called recombination lines. Note that Balmer lines are often
observed in absorption (for example, Balmer lines are evident in a spectrum of the
Sun). This implies that their must be hydrogen present in the first excited state,
which seems at odds with the conclusions reached above. A small fraction of excited
hydrogen atoms, though, can still produce strong absorption lines, simply because
hydrogen is so abundant.

21cm line emission: The ground state of hydrogen is split into two hyperfine
states due to the two possible orientations of the proton and electron spins: The
state in which the spins of proton and electron are aligned has slightly higher energy
than the one in which they are anti-aligned. The energy difference between these

154
two hyperfine states corresponds to a photon with a wavelength of 21cm (which
falls in the radio). The excitation temperature of this spin-flip transition is called
the spin temperature. Since, for typical interstellar medium conditions, the spin-
flip transition is collisionally induced (i.e., the rate for spontaneous de-excitation is
extremely low), the spin-temperature is typically equal to the kinetic temperate of the
hydrogen gas. The 21cm line is an important emission line to probe the distribution
(and temperature) of neutral hydrogen gas in the Universe.

155
CHAPTER 20

The Interaction of Light with Matter:


I - Scattering

To understand radiative processes, and the interaction of photons with matter, it is


important to realize that all photon emission mechanisms arise from accelerating
electrical charge.

The interactions of light with matter can be split in two categories:

• scattering (photon + matter → photon + matter)

• absorption (photon + matter → matter)

We first discuss scattering, which gives rise to a number of astrophysical phenomena:

• reflection nebulae (similar to looking at street-light through fog)

• light echos

• polarization

• Ly-α forest in quasar spectra

The scattering cross-section, σs , is a hypothetical area ([σs ] = cm2 ) which de-


scribes the likelihood of a photon being scattered by a target (typically an electron
or atom). In general, the scattering cross-section is different from the geometrical
cross-section of the particle, and it depends upon the frequency of the photon, and
on the details of the interaction (see below).

Scattering interactions are categorized as either elastic (coherent), where the pho-
ton energy is unchanged by the scattering event, or inelastic (incoherent), where
the photon energy changes.

156
Elastic scattering comes in three forms:

• Thomson scattering γ + e → γ + e

• Resonant scattering γ + X → X + → γ + X

• Rayleigh scattering γ + X → γ + X

Here γ indicates a photon, e a free electron, X an atom or ion, and X + an excited


state of X.

Inelastic scattering comes in two forms:

• Compton scattering γ + e → γ ′ + e′

• Fluorescence γ + X → X ++ → γ ′ + X + → γ ′ + γ ′′ + X

Here accents indicate that the particle has a different energy (i.e., γ ′ is a photon with
a different energy than γ), and X ++ indicates a higher-excited state of X than X + .

In what follows we discuss each of these five processes in more detail.

Thomson scattering: is the elastic (coherent) scattering of electromagnetic radi-


ation by a free charged particle, as described by classical electromagnetism. It is
the low-energy limit of Compton scattering in which the particle kinetic energy
and photon frequency are the same before and after the scattering. In Thomson
scattering the electric field of the incident wave (photon) accelerates the charged
particle, causing it, in turn, to emit radiation at the same frequency as the incident
wave, and thus the wave is scattered. The particle will move in the direction of the
oscillating electric field, resulting in electromagnetic dipole radiation that ap-
pears polarized unless viewed in the forward or backward scattered directions (see
Fig. 26).

The cross-section for Thomson scattering is the Thomson cross section:

8π 2 8πe4
σs = σT = re = 2 4
≃ 6.65 × 10−25 cm2
3 3me c

157
Figure 26: Illustration of how Thomson scattering causes polarization in the direc-
tions perpendicular to that of the incoming EM radiation. The incoming EM wave
~
causes the electron to oscillate in the direction of the oscillation of the E-field. This
acceleration of the electrical charge results in the emission of dipolar EM radiation.

Note that this cross section is independent of wavelength!


In the quantum mechanical view of radiation, electromagnetic waves are made up
of photons which carry both energy (h ν) and momentum (h ν/c). This implies that
during scattering the photon exchanges momentum with the electron, causing the
latter to recoil. This recoil is negligble, until the energy of the incident photon
becomes comparable to the rest-mass energy of the electron, in which case Thomson
scattering becomes Compton scattering.

Compton scattering: is an inelastic scattering of a photon by a free charged par-


ticle, usually an electron. It results in a decrease of the photon’s energy/momentum
(increase in wavelength), called the Compton effect. Part of the energy/momentum
of the photon is transferred to the scattering electron (‘recoil’). In the case of scat-
tering off of electrons at rest, the Compton effect is only important for high-energy
photons with Eγ > me c2 ∼ 0.511 MeV (X-ray and/or gamma ray photons).

158
Figure 27: The Klein-Nishina cross section for Compton scattering. As long as
hν ≪ me c2 one is in the Thomson scattering regime, and σs = σT . However, once
the photon energy becomes comparable to the rest-mass energy of the electron, Comp-
ton scattering takes over, and the cross-section (now called the Klein-Nishina cross-
section), starts to drop as ν −1 .

Because of the recoil effect, the energy of the outgoing photon is



Eγ′ = Eγ
1+ me c2
(1 − cos θ)

where θ is the angle between in incident and outgoing photon. This can also be
written as
λ′ − λ = λC (1 − cos θ)
which expresses that Compton scattering increases the wavelength of the photon by
of order the Compton wavelength λC = h/(me c) ∼ 2.43 × 10−10 cm. If λ ≫ λC
such a shift is negligble, and we are in the regime that is well described by Thomson
scattering.

Compton scattering is a quantum-mechanical process. The quantum aspect also


influences the actual cross section, which changes from the Thomson cross section,
σT , at the low-frequency end, to the Klein-Nishina cross section, σKN (ν), for
hν > me c2 (see Fig. 27). Note how scattering becomes less efficient for more energetic
photons.

159
So far we have considered the scattering of photons off of electrons at rest. A more
realistic treatment takes into account that electrons are also moving, and may do so
relativistically. This adds the possibility of the electron giving some of its kinetic
energy to the photon, which results in Inverse Compton (IC) scattering.

Whether the photon loses (Compton scattering) or gains (IC scattering) energy de-
pends on the energies of the photon and electron. Without derivation, the average
energy change of the photon per Compton scattering against electrons of temperature
Te = me hv 2 i/(3 kB) is
 
∆Eγ 4 kB Te − h ν
=
Eγ me c2

Hence, we have that


h ν > 4 kB Te : Compton effect; photon loses energy to electron
h ν = 4 kB Te : No energy exchange
h ν < 4 kB Te : Inverse Compton effect; electron loses energy to photon

As an example, consider ultra-relativistic electrons with 4 kB Te ≫ h ν. In that


case it can be shown that Inverse Compton (IC) scattering causes the photons to
increase their frequency according to νout ≈ γ 2 νin , where
1
γ=q
v2
1− c2

is the Lorentz factor. Hence, for ultra-relativistic electrons, which have a large
Lorentz factor, the frequency boost of a single IC scattering event can be enormous.
It is believed that this process, upscattering of low energy photons by the IC effect,
is at work in Active Galactic Nuclei.

Another astrophysical example of IC scattering is the Sunyaev-Zel’dovic (SZ)


effect in clusters; the hot (but non-relativistic) electrons of the intra-cluster gas
(with a typical electron temperature of Te ∼ 108 K) upscatter Cosmic Microwave
Background (CMB) photons by a small, but non-negligble amount. The result is a
comptonization of the energy spectrum of the photons; while Compton scattering
maintains photon numbers, it increases their energies, so that they no longer can be
fit by a Planck curve. The strength of this Comptonization (typically expressed in

160
terms of the Compton-y parameter) is measure for the electron pressure Pe ∝ ne Te
along the line-of-sight through the cluster. Observations of the SZ effect provide a
nearly redshift-independent means of detecting galaxy clusters.

Resonant scattering: Resonant scattering, also known as line scattering or


bound-bound scattering is the scattering of photons off electrons bound to nu-
clei in atoms or ions. Before we present the quantum mechanical view of this
process, it is useful to consider the classical one, in which the electron is viewed
as being bound to the nucleus via a spring with a natural, angular frequency,
ω0 = 2πν0 . If the electron is perturbed, it will oscillate at this natural frequency,
which will result in the emission of photons of energy Eγ = hν0 . This in turn
implies energy loss; hence, the bound electron is an example of a damped, har-
monic oscillator. The classical damping constant is given by Γcl = ω02 τe , where
τe = 2e2 /(3me c2 ) ∼ re /c ∼ 6.3 × 10−24 s. This damping, and the corresponding
emission of EM radiation, is the classical analog of spontaneous emission.

Now consider the case of an EM wave of angular frequency, ω, interacting with the
atom/ion. The result is a forced, damped, harmonic oscillator, whose effective
cross section is given by

ω4
σs (ω) = σT
(ω 2 − ω02 )2 + (ω03 τe )2

(see Rybicki & Lightmann 1979 for a derivation). We can distinguish three regimes:

ω ≫ ω0 In this case σs (ω) = σT and we are in the regime of regular Thomson


scattering. The oscillator responds to the high-frequency forcing by adopting the
forced frequency; hence, the system behaves as if the electron is free.

ω ≃ ω0 In this case
σT (Γcl /2)
σs (ω) ≃
2τe (ω − ω0 )2 + (Γcl /2)2
which corresponds to resonant scattering, in which the cross section is hugely
boosted wrt the Thomson case. NOTE: for resonant scattering to be important, it
is crucial that spontaneous de-excitation occurs before collisional excitation or de-
excitation (otherwise the photon energy is lost, and we are in the realm of absorption,
rather than scattering). Typically, this requires sufficiently low densities.

161
Figure 28: Illustration of how Rayleigh scattering causes the sky to be blue. Because
of its strong (λ−4 ) wave-length dependence, blue light is much more scattered than
red light. This causes the Sun light to appear redder than it really is, an effect that
strengthens when the path length through the atmosphere is larger (i.e., at sunrise
and sunset). The blue light is typically scattered multiple times before hitting the
observer, so that it appears to come from random directions on the sky.

ω ≪ ω0 In this case
 4
ω
σs (ω) ≃ σT
ω0
which corresponds to Rayleigh scattering, which is characterized by a strong wave-
length dependence for the effective cross section of the form σs ∝ σT λ−4 .

Rayleigh scattering results from the electric polarizability of the particles. The
oscillating electric field of a light wave acts on the charges within a particle, causing
them to move at the same frequency (recall, the forcing frequency in this case is
much smaller than the natural frequency). The particle therefore becomes a small
radiating dipole whose radiation we see as scattered light.

Rayleigh scattering, and its strong wavelength dependence of σs , is responsible for


the fact that the sky appears blue during the day, and for the fact that sunsets turn
the sky red (see Fig. 28).

162
We now turn our attention to a quantum-mechanical view of resonant scatter-
ing. The main difference between the classical view (above) and the quantum view
(below), is that in the latter there is not one, but many ‘natural frequencies’, νij , cor-
responding to all the possible energy-level-transitions ∆Eij = hνij that correspond
to the atom/ion in question.

Oscillator strength: With each transition corresponds an oscillator strength, fij ,


which is a dimensionless quantity that expresses the ‘strength’ of the i ↔ j transition.
It expresses the quantum mechanical probability that transition i → j occurs under
the incidence of a νij photon given the quantum-mechanical selection rules, which
state the degree to which a certain transition between degrees is allowed. You can
think of fij as being proportional to the probability that the incidence of a νij photon
results in the corresponding electronic transition.

In the quantum-mechanical view, the three regimes of bound-bound scattering have


effective cross sections:

σs = σT ν ≫ νij Thomson scattering


π e2
σs (ν) = fij φL (ν)
me c
ν ≃ νij Resonant scattering
 4
σs (ν) = σT fij ννij ν ≪ νij Rayleigh scattering

Here φL (ν) is the Lorentz profile, which describes the natural line broadening
associated with the transition in question. The non-zero width of this Lorentz profile
implies that resonant scattering is not perfectly coherent; typically the energy of the
outgoing photon will be slightly different from that of the incident photon. The
probability distribution for this energy shift is described by φL (ν), and originates
from the Heisenberg Uncertainty Principle, according to which ∆E ∆t ≥ h̄/2;
hence, the uncertainty related to the time it takes for the electron to spontaneously
de-excite results in a related ‘uncertainty’ in energy.

Fig. 29 shows the frequency dependence of a (quantum-mechanical) atom/ion. It


shows the Rayleigh regime at small ν, the resonant scattering peaks at a few
transition frequencies, and the Thomson regime at large ν. Note that the height
of the various peaks are set by their respective oscillator strengths.

163
Figure 29: Illustration of the scattering cross section of an atom or ion with at
least one bound electron. At high (low) frequency, scattering is in the Thomson
(Rayleigh) regime; at specific, intermediate frequencies, set by the transition energies
of the atom/ion, resonant scattering dominates; the profiles are Lorentz profiles, and
reflect the natural line broadening. The relative heights of the peaks are set by their
oscillator strengths. NOTE: figure is not to scale; typically the cross section for
resonant scattering is orders of magnitude larger than the Thomson cross section.

164
Figure 30: Example of a quasar spectrum revealing the Ly-α forest due to resonant
scattering of Ly-α photons by neutral hydrogen along the line-of-sight from quasar to
observer.

An important, astrophysical example of resonant scattering is the Ly-α forest in the


spectra of (high-redshift) quasars. Redward of the quasar’s Ly-α emission line one
typically observes a ‘forest’ of ‘absorption lines’, called the Ly-α forest (see Fig. 30).
These arise from resonant scattering in the Ly-α line of neutral hydrogen in gas clouds
along the line-of-sight between the quasar and observer. NOTE: although these are
called ‘absorption lines’ they really are a manifestation of (resonant) scattering.

Fluorescence: fluorescence is an inelastic (incoherent) scattering mechanism, in


which a photon excites an electron by at least two energy states, and the spontaneous
de-excitation occurs to one or more of the intermediate energy levels. Consequently,
the photon that is ‘scattered’ (i.e., absorbed and re-emitted) has changed its energy.

165
CHAPTER 21

The Interaction of Light with Matter:


II - Absorption

Absorption: The absorption of photons can have three effects:

• heating of the absorbing medium (heating of dust grains, or excitation of


gas followed by collisional de-excitation)

• acceleration of absorbing medium (radiation pressure)

• change of state of absorbing medium (ionization, sublimation or dissoci-


ation)

Note that ionization (transition from neutral to ionized), sublimation (transition


from solid to gas) and dissociation (transition from molecular to atomic) can also
occur as a consequence of particle collisions. Therefore one often uses terms such as
photo-ionization and collisional ionization to distinguish between these.

Photoionization: Photoionization is the process in which an atom is ionized by


the absorption of a photon. For hydrogen, this is

HI + γ → p + e ,

where HI denotes a neutral hydrogen atom. Using that the rate, Γ, of an interaction
is always given by Γ = t−1
coll = nσv, the photoionization rate, Γγ,H , is proportional
to the number density of ionizing photons and to the photoionization cross section,
σpi (ν), according to:
Z ∞
Γγ,H = c σpi (ν) nγ (ν) dν
νt

where νt is the threshold frequency for ionization (corresponding to 13.6eV in the


case of hydrogen). nγ (ν)dν is the number density of photons with frequencies in the
range ν to ν + dν, and is related to the energy flux of the radiation field, J(ν), by

166
4 π J(ν)
nγ (ν) = .
chν
The photoionization cross sections can be obtained from quantum electrodynamics
by calculating the bound-free transition probability of an atom in a radiation field
(see e.g., Rybicki & Lightman 1979).

Recombination: Recombination is the process by which an ion recombines with


an electron. For hydrogen ions (i.e. protons), the process is

p + e → HI + γ .

For hydrogen (or a hydrogenic ion, i.e., an ion with a single electron), the recom-
bination cross section to form an atom (or ion) at level n, σrec (v, n), is related to
the corresponding photoionization cross section by the Milne relation:

 2
gn hν
σrec (v, n) = σpi (ν, n) ,
gn+1 me c v

where gn = 2n2 is the statistical weight of energy level n and ν and v are related by
me v 2 /2 = h(ν − νn ), with hνn the threshold energy required to ionize an atom whose
electron sits in energy state n. The recombination coefficient for a given level n is
the product of the capture cross section and velocity, σrec (v, n) v, averaged over the
velocity distribution f (v). For an optically thin gas where all photons produced by
recombination can escape without being absorbed, the total recombination coefficient
is the sum over all n:


X ∞ Z
X
αA = αn = σrec (v, n) v f (v) dv
n=1 n=1

This is called the Case A recombination coefficient, to distinguish it from the


Case B recombination in an optically thick gas. In Case B, recombinations to the
ground level generate ionizing photons that are absorbed by the gas, so that they
do not contribute to the overall ionization state of the gas. It is easy to see that the
Case B recombination coefficient is αB = αA − α1 .

167
Strömgren sphere: A sphere of ionized hydrogen (H II) around an ionizing source
(e.g., AGN, O or B star, etc.). Ionization of hydrogen (from the ground state) requires
a photon energy of at least 13.6eV, which implies UV photons. In a (partially)
ionized medium, electrons and nuclei recombine to produce neutral atoms. The
region around an ionizing source will ultimately establish ionization equilibrium
in which the number of ionizations is equal to the number of recombinations.

Consider an ionizing source in a uniform medium of pure hydrogen. Let Ṅion be the
number of ionizing photons produced per second. The corresponding recombination
rate is given by
4
Ṅrec = ne np αrec V = n2e αB πRs3
3
where we have used that, for a pure hydrogen gas, ne = np , and Rs is the radius of
the Strömgren sphere (i.e., the radius of the sphere that is going to be ionized),
which can be written as !1/3
3 Ṅion
Rs =
4 π αB n2e

Using that the luminosity of the ionizing source, L∗ , is related to its surface intensity,
I∗ , according to
L∗ = 4 π R∗2 F∗ = 4 π 2 R∗2 I∗
where R∗ is the radius of the ionizing source (i.e., an O-star) and we have used that
F∗ = π I∗ (see Chapter 18). Hence, we have that
Z ∞ Z ∞
2 2 Bν (T ) π L∗ Bν (T )
Ṅion = 4 π R∗ dν = 4

νt hν σSB Teff νt hν

where we have assumed that the ionizing source is a Black Body of temperature T ,
and, in the second part, that L∗ = 4πR∗2 σSB Teff
4
.

Thus, by measuring the luminosity and effective temperature of a star, and the radius
of its Strömgren sphere, one can infer the (electron) density of its surroundings.

168
CHAPTER 22

The Interaction of Light with Matter:


III - Extinction

As we have seen, there are numerous processes by which a photon can interact with
matter. It is useful to define the mean-free path, l, for a photon and the related
opacity and optical depth.

Opacity: a measure for the impenetrability to electro-magnetic radiation due to the


combined effect of scattering and absorption. If the opacity is caused by dust we call
it extinction.

Optical Depth: the dimensionless parameter, τν , describing the opacity/extinction


at frequency ν. In particular, the infinitesimal increase in optical depth along a line
of sight, dτν , is related to the infinitesimal path length dl according to

dτν = σν n dl = κν ρ dl = αν dl

Here σν is the effective cross section ([σν ] = cm2 ), κν is the mass absorption
coefficient ([κν ] = cm2 g−1 ), αν is the absorption coefficient ([αν ] = cm−1 ), and
n and ρ are the number and mass densities, respectively. The optical depth to a
source at distance d is therefore
Z d Z d
τν = dτν = κν (l) ρ(l) dl
0 0

The ISM/IGM between source and observer is said to be optically thick (thin) if
τν > 1 (τν < 1).

Opacity/extinction reduces the intensity of a source according to

Iν,obs = Iν,0 e−τν

where Iν,0 is the unextincted intensity (i.e., for τν = 0).

169
Rosseland Mean Opacities: In the case of stars opacity is crucially important
for understanding stellar structure. Opacities within stars are typically expressed
in terms of the Rosseland mean opacities, κ, which is a weighted average of κν
over frequency. Typically, one finds that κ ∝ ρ T −3.5 , which is known as Kramer’s
opacity law, and is a consequence of the fact that the opacity is dominated by
bound-free and/or free-free absorption. A larger opacity implies stronger radiation
pressure, which gives rise to the concept of the Eddington luminosity.

Eddington Luminosity: the maximum luminosity a star (or, more general, emit-
ter) can achieve before the star’s radiation pressure starts to exceed the force of
gravity.

Consider an outer layer of a star with a thickness l such that τν ≃ κν ρl = 1. Then, a


photon passing this layer will be absorbed and contribute to radiation pressure. The
resulting force exerted on the matter is
L
Frad =
c
This has to be compared to the gravitational force
G M∗ mlayer
Fgrav =
r2
Using that mlayer = 4πr 2lρ = 4πr 2 /κν , we find that Frad = Fgrav if the luminosity is
equal to

4 π G M∗ c
LEdd =
κν

which is called the Eddington luminosity. Stars with L > LEdd cannot exist, as
they would blow themselves appart (Frad > Fgrav ). The most massive stars known
have luminosities that are very close to their Eddington luminosity.

Since the luminosities of AGN (supermassive black holes with accretion disks) are set
by their accretion rate, the same argument implies an upper limit to the accretion
rate of AGN, known as the Eddington limit.

170
Figure 31: Empirical extinction laws, defined in terms of the ratio of color excesses,
for the Milky Way (MW), the Large Magellanic Cloud (LMC) and the Small Mag-
ellanic Cloud (SMC). Note the strong feature around 2100 Å in the MW extinction
law, believed to be due to graphite dust grains.

Extinction by Dust: Dust grains can scatter and absorb photons. Their ability
to do so depends on (i) grain size, (ii) grain composition, and (iii) the presence of a
magnetic field, which can cause grain alignment. Observationally, the extinction in
the V -band is defined by
   
fV IV
AV ≡ −2.5 log = −2.5 log
fV,0 IV,0

where the subscript zero refers to the unextincted flux/intensity. Using that IV =
IV,0 e−τV we have that
AV = 1.086 τV
More generally, Aλ = 1.086 τλ ; hence, an optical depth of unity roughly corresponds
to an extinction of one magnitude.

Reddening: In addition to extinction, dust also causes reddening, due to the fact
that dust extinction is more effective at shorter (bluer) wavelengths.

171
Color Excess: E(B − V ) ≡ AB − AV , which can also be defined for any other
wavebands.

Extinction law: conventionally, dust extinction is expressed in terms of an empirical


extinction law:
Aλ Aλ
k(λ) ≡ ≡ RV
E(B − V ) AV
where
AV
RV ≡
E(B − V )
is a quantity that is insensitive to the total amount of extinction; rather it expresses
a property of the extinction law. Empirically, dust in the Milky Way seems to have
RV ≃ 3.1, while the dust in the Small Magellanic Cloud (SMC) is better characterized
by RV ≃ 2.7. Dust extinction law is not universal; rather, it is believed to depend
on the local UV flux and the metallicity, among others (see Fig. 31).

Theoretical attempts to model the extinction curve have shown that dust comes in
two varieties, graphites and silicates, while the grain-size distribution is well fit by
dN/da ∝ a−3.5 and covers the range from ∼ 0.005µm to ∼ 0.25µm. Note that for
radiation with λ > amax ≃ 2500 Å dust mainly causes Rayleigh scattering.

172
CHAPTER 23

Radiative Transfer

Consider an incoming signal of specific intensity Iν,0 passing through a cloud (i.e.,
any gaseous region). As the radiation transits a small path length dr through the
cloud, its specific intensity changes by dIν = dIν,loss +dIν,gain . The loss-term describes
the combined effect of scattering and absorption, which remove photons from the line-
of-sight, while the gain-term describes all processes that add photons to the line-of-
sight; these include all emission processes from the gas itself, as well as scattering of
photons from any direction into the line-of-sight.

In what follows we ignore the contribution of scattering to dIν,gain , as this term makes
solving the equation of radiative transfer much more complicated. We will briefly
comments on that below, but for now the only process that is assumed to contribute
to dIν,gain are emission processes from the gas.

It is useful to define the following two coefficients:

• Absorption coefficient, αν = n σν = ρ κν , which has units [αν ] = cm−1 .

• Emission coefficient, jν , defined as the energy emitted per unit time, per unit
volume, per unit frequency, per unit solid angle (i.e., dE = jν dt dV dν dΩ, and thus
[jν ] = erg s−1 cm−3 Hz−1 sr−1 ).

In terms of these two coefficients, the equation of radiative transfer can be


written in either of the following two forms:

dIν
= −αν Iν + jν (form I)
dr
dIν
= −Iν + Sν (form II)
dτν

173
Here Sν ≡ jν /αν is called the source function, and has units of specific intensity
(i.e., [Sν ] = erg s−1 cm−2 Hz−1 sr−1 ). In order to derive form II from form I, recall
that dτν = αν dr (see Chapter 22).

NOTE: we use the convention of τν increasing from the source towards the observer.
Some textbooks adopt the opposite convention, which results in some sign differences.

To get some insight, we now consider a number of different cases:

Case A No Cloud
In this case, there is no absorption (αν = 0) or emission (jν = 0), other than the
emission from the background source. Hence, we have that
dIν
=0 ⇒ Iν = Iν,0
dr
which expresses that intensity is a conserved quantity in vacuum.

Case B Absorption Only


In this case, the cloud absorbs background radiation, but does not emit anything
(jν = Sν = 0). Hence,
dIν
= −Iν
dτν
which is easily solved to yield
Iν = Iν,0 e−τν
which is the expected result (see Chapter 22).

Case C Emission Only


If the cloud does not absorb (αν = 0) but does emit we have
Z l
dIν
= jν ⇒ Iν = Iν,0 + jν (r) dr
dr 0

where l is the size of the cloud along the line-of-sight. This equation simply expresses
that the increase of intensity is equal to the emission coefficient integrated along the
line-of-sight.

174
Case D Cloud in Thermodynamic Equilibrium w/o Background Source

Consider a cloud in TE, i.e., specified by a single temperature T (kinetic temperature


is equal to radiation temperature). Since in a system in TE there can be no net
transport of energy, we have that
dIν jν
= −αν Iν + jν = 0 ⇒ Iν = = Sν
dr αν
Since the observer must see a black body of temperature T , we also have that Iν =
Bν (T ) (i.e., the intensity is given by a Planck curve corresponding to the temperature
of the cloud), and we thus have that

Iν = Sν = Bν (T )
jν = αν Bν (T )

The latter of these equivalent relations is sometimes called Kirchoff’s law, and
simply expresses that a black body needs to establish a balance between emission
and absorption (i.e., Bν (T ) = jν /αν ).

Case E Emission & Absorption (formal solution)


Consider the general case with both emission and absorption (but where we ignore
the fact that scattering can scatter photons into the line of sight). Starting from form
II of the equation of radiative transfer, multiplying both sides with eτν , we obtain
that
dI˜ν
= −S̃ν
dτν
where I˜ν ≡ Iν eτν and S̃ν ≡ Sν eτν . We can rewrite the above differential equation as
Z I˜ν Z τν
dI˜ν = S̃ν dτν′
I˜ν,0 0

Using that I˜ν,0 = Iν,0 e0 = Iν,0 the solution to this simple integral equation is
Z τν
−τν ′
Iν = Iν,0 e + Sν (τν′ ) e−(τν −τν ) dτν′
0

175
where τν is the total optical depth along the line of sight (i.e., through the cloud).
The above is the formal solution, which, under the simplifying assumption that the
source function is constant along the line of sight reduces to

Iν = Iν,0 e−τν + Sν 1 − e−τν

The first term expresses the attenuation of the background signal, the second
term expresses the added signal due to the emission from the cloud, while
the third term describes the cloud’s self-absorption.

Using the above formal solution to the equation of radiative transfer, we have the
following two extremes:

τν ≫ 1 ⇒ I ν = Sν
τν ≪ 1 ⇒ Iν = Iν,0 (1 − τν ) + Sν τν

where, for the latter case, we have used the Taylor series expansion for the exponen-
tial. In the high optical depth case, the observer just ‘sees’ the outer layers of the
cloud, and therefore the observed intensity is simply the source function of the cloud
(the observed signal contains no contribution from the background source). In the
small optical depth limit, the contribution from the cloud is suppressed by a factor
τν , while that from the background source is attenuated by a factor (1 − τν ). Using
that Sν = jν /αν and τν = αν l (if the absorption coefficient is constant throughout the
cloud), we see that τν Sν = jν l; in other words, the contribution from the cloud itself
is simply its emission coefficient (assumed constant throughout the cloud) multiplied
with the pathlength through the cloud.

To get some further insight into the source function and radiative transfer in general,
consider form II of the radiate transfer equation. If Iν > Sν then dIν /dτν < 0, so that
the specific intensity decreases along the line of sight. If, on the other hand, Iν < Sν
then dIν /dτν > 0, indicating that the specific intensity increases along the line of
sight. Hence, Iν tends towards Sν . If the optical depth of the cloud is sufficiently
large than this ‘tendency’ will succeed, and Iν = Sν .

An important special case of the general solution derived above is if the cloud is in
local thermal equilibrium (LTE). This is very often the case, since over the mean
free path of the photons, every system will tend to be in LTE, unless it was recently

176
disturbed and has not yet been able to equilibrate. In the case of LTE, we have that,
over a patch smaller than or equal to the mean free path of the photons, we have that
Sν ≡ jν /αν = Bν (T ), where T is the kinetic temperature (= radiation temperature)
of the patch.

The solution to the equation of radiative transfer now is


 
Iν = Iν,0 e−τν + Bν (T ) 1 − e−τν

Note that Iν is not constant throughout the cloud, as was the case for a cloud in
TE. In the case of LTE, however, there can be a non-zero gradient dIν /dr.

Keeping this difference in mind, we now look at the solution to our equation of
radiative transfer for a cloud in LTE at its two extremes:
τν ≫ 1 ⇒ Iν = Bν (T )
τν ≪ 1 ⇒ Iν = Iν,0 (1 − τν ) + Bν (T ) τν
The former expresses that an optically thick cloud in LTE emits black body radiation.
This is characterized by the fact that (i) if there is a background source, you can’t see
it, (ii) you can look into the source only for about one mean free path of the photons
(which is much smaller than the size of the source), and (iii) the only information
available to an observer is the temperature of the cloud (the observed intensity is a
Planck curve of temperature T ).

A good example of gas clouds in LTE are stars!

In the optically thin limit, the observed intensity depends on the background source
(if present), and depends on both the temperature (sets source function) and density
(sets optical depth) of the cloud (recall that τν ∝ κν ρ l).

In the case without background source we have that



Bν (T ) if τν ≫ 1
Iν =
τν Bν (T ) if τν ≪ 1
Note that this is different from case D, in which we considered a cloud in TE without
background source. In that case we obtained that Iν = Bν (T ) independent of τν .

177
In the case of LTE, however, there are radial gradients, which are responsible for
diminishing the intensity by the optical depth in the case where τν ≪ 1. This
may seem somewhat ‘counter-intuitive’, as it indicates that a cloud of larger optical
depth is more intense!!! To understand this, consider the limit τν → 0. In this case
all photons pass through the cloud with zero probability to be absorbed/scattered.
In this situation, there is simply no way to establish an equilibrium between emission
and absorption required for the establishment of a black body; or, put differently,
if there is no absorption, there is no emission either (after all, we are in LTE), and
thus, Iν = 0.

Based on the above, we have that, in the case of a cloud in LTE without background
source, Iν ≤ Bν (T ), where T is the temperature of the cloud. If we express the
intensity in terms of the brightness temperature we have that TB,ν ≤ T . Hence,
for a cloud in LTE without background source the observed brightness temperature
is a lower limit on the kinetic temperature of the cloud.

What about scattering? In the most general case, any element in the cloud re-
ceives radiation coming from all 4π sterradian, and a certain fraction of that radiation
will be scattered into the line-of-sight of an observer.

In general, the scattering can (will) be non-isotropic (e.g., Thomson scattering) and
incoherent (e.g., Compton scattering or resonant scattering), and the final equation
of radiative transfer can only be solved numerically.

In the simplified case of isotropic, coherent scattering the corresponding emission


coeffient can be found by simply equating the power absorbed per unit volume to
that emitted (for each frequency);

jν,scat = αν,scat Jν

where αν,scat is the absorption coefficient of the scattering processes, while


Z
1
Jν = Iν dΩ

is the mean intensity, averaged over all 4π sterradian.

178
The source function due to scattering is then simply
Z
jν,scat 1
Sν ≡ = Jν = Iν dΩ
αν,scat 4π

Hence, the source function due to isotropic, coherent scattering is simply the
mean intensity.

The radiative transfer equation for pure scattering (no background source, and no
emission) is
dIν
= −αν,scat (Iν − Jν )
dr
Even this oversimplified case of pure isotropic, coherent scattering is not easily solved.
Since Jν involves an integration (over all 4π sterradian), the above equation is an
integro-differential equation, which are extremely difficult to solve in general; one
typically has to resort to numerical methods (see Rybicki & Lightmann 1979 for
more details). NOTE: although the scattering may be isotropic, the incoming radia-
tion is typically not.

Observability of Emission & Absorption Lines: Consider a cloud in front of


some background source. Assume the cloud is in LTE at temperature T . Assume
that αν is only non-zero at a specific frequency, ν1 , corresponding to some electron
transition. Given that resonant scattering is typically orders of magnitude more
efficient than other scattering mechanisms, this is a reasonable approximation. The
intensity observed is  
Iν = Iν,0 e−τν + Bν (T ) 1 − e−τν
At all frequencies other than ν1 we have τν = 0, and thus Iν = Iν,0 . Now assume
that the observer sees an absorption line at ν = ν1 . This implies that
 
Iν1 = Iν1 ,0 e−τν1 + Bν1 (T ) 1 − e−τν1 < Iν1 ,0

while in the case of an emission line


 
Iν1 = Iν1 ,0 e−τν1 + Bν1 (T ) 1 − e−τν1 > Iν1 ,0

Rearranging, we then have that

179
Absorption Line: Bν (T ) < Iν,0 T < TB,ν

Emission Line: Bν (T ) > Iν,0 T > TB,ν

where T is the (kinetic) temperature of the cloud, and TB,ν is the brightness tem-
perature of the background source, at the frequency of the line. Hence, if the cloud
is colder (hotter) than the source, an absorption (emission) line will arise. In the
case of no background source, we effectively have that TB,ν = 0, and the cloud will
thus reveal an emission line. In the case where T = TB,ν no line will be visible,
independent of the optical depth of the cloud!

180
CHAPTER 24

Continuum Emission Mechanisms

Continuum radiation is any radiation that forms a continuous spectrum and is not
restricted to a narrow frequency range. In what follows we briefly describe five
continuum emission mechanisms:

• Thermal (Black Body) Radiation

• Bremsstrahlung (free-free emission)

• Recombination (free-bound emission)

• Two-Photon emission

• Synchrotron emission

In general, the way to proceed is to ‘derive’ the emission coeffient, jν , the absorption
coefficient, αν , and then use the equation of radiative transfer to compute the
specific intensity, Iν , (i.e., the ‘spectrum’), for a cloud of gas emitting continuum
radiation using any one of those mechanisms.

First some general remarks: when talking about continuum processes it is important
to distinguish thermal emission, in which the radiation is generated by the thermal
motion of charged particles and in which the intensity therefore depends (at least) on
temperature, i.e., Iν = Iν (T, ..), from non-thermal emission, which is everything
else.

Examples of thermal continuum emission are black body radiation and (thermal)
bremsstrahlung, while synchrotron radiation is an example of non-thermal emission.
Another non-thermal continuum mechanism is inverse compton radiation. However,
since this is basically an incoherent photon-scattering mechanism, rather than a
photon-production mechanism, we will not discuss IC scattering any further here.

181
Characteristics of Thermal Continuum Emission:

• Low Brightness Temperatures: Since one rarely encouters gases with kinetic
temperatures T > 107 − 108 K, and since TB ≤ T (see Chapter 28), if the brightness
temperature of the radiation exceeds ∼ 108 K it is most likely non-thermal in origin
(or has experienced IC scattering).

• No Polarization: Since these is no particular directionality to the thermal motion


of particles, thermal emission is essentially unpolarized. In other words, if emission
is found to be polarized, it is either non-thermal, or the signal became polarized after
it was emitted (i.e., via Thomson scattering).

Thermal Radiation & Black Body Radiation: Thermal radiation is the con-
tinuum emission arising from particles colliding, which causes acceleration of charges
(atoms typically have electric or magnetic dipole moments, and colliding those results
in the emission of photons). This thermal radiation tries to establish thermal equi-
librium with the matter that produces it via photon-matter interactions. If thermal
equilibrium is established (locally), then the source function Sν ≡ jν /αν = Bν (T )
(Kirchoff’s law).
As we have seen in the previous Chapter;

Bν (T ) if τν ≫ 1
Iν =
τν Bν (T ) if τν ≪ 1

where τν = αν l is the optical depth through the cloud, which has a dimension l
along the line-of-sight.

Free-free emission (Bremsstrahlung): Bremsstrahlung (German for ‘braking


radiation’) arises when a charged particle (i.e., an electron) is accelerated though
the Coulomb interaction with another charged particle (i.e., an ion of charge Ze).
Effectively what happens is that the two charges make up an electric dipole which,
due to the motion of the charges, is time variable. A variable dipole is basically an
antenna, and emits electromagnetic waves. The energy in these EM waves (photons)
emitted is lost to the electron, which therefore loses (kinetic) energy (the electron is
‘braking’).

It is fairly straightforward to compute the amount of energy radiated by a single

182
electron moving with velocity v when experiencing a Coulomb interaction with a
charge Ze over an impact parameter b (see Rybicki & Lightmann 1979 for a de-
tailed derivation).

The next step is to integrate over all possible impact parameters. This are all impact
parameters b > bmin , where from a classical perspective bmin is set by the requirement
that the kinetic energy of the electron, Ek = 12 me v 2 , is larger than the binding
energy, Eb = Ze2 /b (otherwise we are in the regime of recombination; see below).
However, there are some quantum mechanical corrections one needs to make to
this bmin which arise from Heisenberg’s Uncertainty Principle (∆x ∆p ≥ h̄/2). This
correction factor is called the free-free Gaunt factor, gff (ν, Te ), which is close to
unity, and has only a weak frequency dependence. The final step in obtaining the
emission coeffient is the integration over the Maxwellian velocity distribution of
the electrons, characterized by Te . The result (in erg s−1 cm−3 Hz−1 sr−1 ) is:

 
−39 Z2
jν = 5.44 × 10 1/2
ne ni gff (ν, Te ) e−hν/kB Te
Te
In the case of a pure (ionized) hydrogen gas, Z = 1 and ni = ne . Upon inspection, it
is clear that free-free emission has a flat spectrum jν ∝ ν α with α ∼ 0 (controlled by
the weak frequency dependence of the Gaunt factor) with an exponential cut-off for
h ν > kB Te (the maximum photon energy is set by the temperature of the electrons).
This reveals that a measurement of the exponential cut-off is a direct measure of the
electron temperature.

The above emission coefficient tells us the emissive behavior of a pocket of (ion-
ized) gas without allowance for the internal absorption. Accounting for the latter
requires radiative transfer. Since Bremsstrahlung arises from collisions, we may use
the LTE approximation. Hence, Kirchoff’s law tells us that αν = jν /Bν (T ), which
allows us to compute the absorption coefficent, and thus the optical depth τν = αν l.
Substitution of Bν (T ), with T = Te , yields
τν ≃ 3.7 × 108 Z 2 Te−1/2 ν −3 [1 − e−hν/kB Te ] gff (ν, Te ) E
where Z
E≡ n2e dl ≃ n2e l
is called the emission measure, and we have assumed that ne = ni . Upon in-
spection, one notices that τν ∝ ν −2 (for hν ≪ kB Te ), indicating that the opacity

183
of the cloud increases with decreasing frequency. This opacity arises from free-free
absorption, which is simply the inverse process of free-free emission; a photon is
absorbed by an electron that is experiencing a Coulomb interaction.

If we now substitute our results in the equation of radiative transfer (without


background source),  
Iν = Bν (T ) 1 − e−τν
then we obtain that 
Bν (Te ) if τν ≫ 1
Iν =
τν Bν (T ) = jν l if τν ≪ 1

Fig. 32 shows an illustration of a typical free-free emission spectrum: at low frequency


the gas is optically thick, and one probes the Rayleigh-Jeans part of the Planck curve
corresponding to the electron temperature (Iν ∝ ν 2 Te ). At intermediate frequencies,
−1/2
where the cloud is optically thin, the spectrum is flat (Iν ∝ ETe ), and at the
high-frequency end there is an exponential cut-off (Iν ∝ exp[−hν/kB Te ]).

Free-Bound emission (Recombination): this involves the capture of a free elec-


tron by a nucleus into a quantized bound state. Hence, this requires the medium to
be ionized, similar to free-free emission, and in general both will occur (complicating
the picture). Free-bound emission is basically the same as free-free emission (they
have the same emission coefficient, jν ), except that they involve different integra-
tion ranges for the impact parameter b, and therefore different Gaunt factors; the
free-bound Gaunt factor gfb (ν, Te ) has a different temperature dependence than
gff (ν, Te ), and also has more ‘structure’ in its frequency dependence; in the limit where
the bound state has a large quantum number (i.e., the electron is weakly bound), we
have that gfb ∼ gff . However, for more bound states the frequency dependence of gfb
reveals sharp ‘edges’ associated with the discrete bound states.

When kB Te ≫ h ν recombination is negligible (electrons are moving too fast to be-


come bound), and the emission is dominated by the free-free process (i.e., gfb(ν, Te ) →
0 for hν ≪ kB Te ) At lower electron temperatures (or, equivalently, higher photon
frequencies), recombination becomes more and more important, and often will dom-
inate over bremsstrahlung, gfb (ν, Te ) > gff (ν, Te ), (see Fig. 33).

184
Figure 32: Specific intensity of free-free emission (Bremssstrahlung), including the
effect of free-free self absorption at low frequencies, where the optical depth exceeds
unity. At low frequencies, one probes the Rayleigh-Jeans part of the Planck curve
corresponding to the electron temperature. At intermediate frequencies, where the
cloud is optically thin, the spectrum is flat, followed by an exponential cut-off at the
high-frequency end.

185
Two-Photon Emission: two photon emission occurs between bound states in an
atom, but it produces continuum emission rather than line emission.

Two photon emission occurs when an electron finds itself in a quantum level for which
any downward transition would violate quantum mechanical selection rules. Each
transition is therefore highly forbidden. However, there is a non-zero chance that
the electron decays under the emission of two, rather than one, photons. Energy
conservation guarantees that ν1 + ν2 = νtr = ∆Etr /h, where ∆Etr is the energy
difference associated with the transition. The most probable configuration is the
one in which ν1 = ν2 = νtr /2, but all configurations that satisfy the above energy
conservation are possible; they become less likely the larger |ν1 − νtr /2|, resulting in
a ‘continuum’ emission that appears as an extremely broad ‘emission line’. In fact,
whereas the number of photons with 0 < ν < νtr/2 is equal to that with νtr/2 < ν < νtr ,
the latter have more energy (i.e., Eγ = hν). Consequently, the spectral energy
distribution, Lν (erg s−1 Hz−1 ) is skewed towards higher frequency.

For two photon emission to occur, we require that spontaneous emission happens
before collisional de-excitation has a chance. Consequently, two-photon emission oc-
curs in low density ionized gas. The strength of the two photon emission depends on
the number of particles in the excited states. This in turn depends on the recombi-
nation rate; although two-photon emission is quantum-mechanical in nature, it can
still be throught of as ‘thermal emission’, and the density dependence is the same as
for free-free and free-bound emission (i.e., jν ∝ n2e ).

An important example of two-photon emission is associated with the Lyα recom-


bination line, which results from a de-excitation of an electron from the n = 2 to
n = 1 energy level in a Hydrogen atom. As it turns out, the n = 2 quantum level
consists of both 2s and 2p states. The transition 2p → 1s is a permitted transition
with A2p→1s = 6.27 × 108 s−1 . However, the 2s → 1s transition is highly forbidden,
and has a two-photon-emission rate coefficient of A2s→1s = 8.2s−1 . Although this is
orders of magnitude lower than for the 2p → 1s transition, the two photon emission
< 104 cm−3 ).
can still be important in low-density nebulae (n ∼

Synchrotron & Cyclotron Emission: A free electron moving in a magnetic field

186
Figure 33: Emission spectra of plasmas with solar abundances. The histogram indi-
cates the total spectrum, including line radiation. The spectrum has been binned in
order to highlight the relative importance of line radiation. The thick solid line is the
total continuum emission, the thin solid line the contribution due to Bremsstrahlung,
the dashed line free-bound emission and the dotted line two-photon emission. Note
how recombination becomes less and less important when the gas gets hotter. [From
Kaastra et al. 2008, Space Science Reviews, 134, 155]

187
experiences a Lorentz force:
 
~
v ~ = e v B sin φ = e v B⊥ = e v⊥ B
F~e = e ×B
c c c c

where φ is the pitch angle between ~v and B.~ If φ = 0 the particle moves along the
magnetic field, and the Lorentz force is zero. If φ = 90o the particle will move in a
circle around the magnetic field line, while for 0o < φ < 90o the electron will spiral
(‘cork-screw’) around the magnetic field line. In the latter two cases, the electron is
being accelerated, which causes the emission of photons. Note that this applies to
both electrons and ions. However, since the cyclotron (synchrotron) emission from
ions is negligble compared to that from electrons, we will focus on the latter.

If the particle is non-relativistic, then the emission is called cyclotron emission. If,
on the other hand, the particles are relativistic, the emission is called synchrotron
emission. We will first focus on the former.

Cyclotron emission: the gyrating electron emits dipolar emission that (i) has the
frequency of gyration, and (ii) is highly polarized. Depending on the viewing angle
~ linear
the observer can see circular polarization (if line-of-sight is alined with B),
~ or elliptical polarization (for any
polarization, if line of sight is perpendicular to B,
other orientation).

The gyration frequency can be obtained by equating the Lorentz force with the
centripetal force:
2
e v⊥ me v⊥
Fe = B=
c r0
where v⊥ = v sin φ, which results in
me v⊥ c
r0 =
eB
~ This is called the gyration radius (or gyro-radius). The period of
where B = |B|.
gyration is T = 2πr0 /v⊥ , which implies a gyration frequency (i.e., the frequency
of the emitted photons) of
1 eB
ν0 = =
T 2π me c

188
Note that this frequency is independent of the velocity of the electron! It only depends
on the magnetic field strength B;

ν0 ~
|B|
= 2.8
MHz Gauss
We thus see that cyclotron emission really is line emission, rather than continuum
emission. The nature of this line emission is very different though, from ‘normal’
spectral lines which result from quantum transitions within atoms or molecules.
Note, though, that if the ‘source’ has a varying magnetic field, then the variance in
B will result in a ‘broadening’ of the line, which, if sufficiently large, may appear as
‘continuum emission’.

In principle, observing cyclotron emission immediately yields the magnetic field


strength. However, unless B is extremely large, the frequency of the cyclotron emis-
sion is extremely low; typical magnetic field strengths in the IGM are of the order of
several µG, which implies cyclotron frequencies in the few Hz regime. The problem is
that such low frequency radiation will not be able to travel through an astrophysical
plasma, because the frequency is lower than the plasma frequency, which is the
natural frequency of a plasma (see Chapter 20). In addition, the Earth’s ionosphere
< 10MHz, so that we can only observe cyclotron
blocks radiation with a frequency ν ∼
emission from the Earth’s surface if it originates from objects with B ∼ > 3.5G. For

this reason, cyclotron emission is rarely observed, with the exception of the Sun,
some of the planets in our Solar System, and an occasional pulsar.

Synchrotron Emission: this is the same as cyclotron emission, but in the limit in
which the electrons are relativistic. As we demonstrate below, this has two important
effects: it makes the gyration frequency dependent on the energy (velocity) of the
electron, and it causes strong beaming of the electron’s dipole emission.

In the relativistic regime, the electron energy becomes Ee = γ me c2 , where γ =


(1 − v 2 /c2 )−1/2 is the Lorentz factor. This boosts the gyration radius by a factor
γ, and reduces the gyration frequency by 1/γ:
γ me v⊥ c γ me c2
r0 = ≃
eB eB
eB
ν0 =
2π γ me c

189
Figure 34: Illustration of how the Lorentz transformation from the electron rest frame
to the lab frame introduce relativistic beaming with an opening angle θ = 1/γ. Note
that in the electron rest frame, the synchrotron emission is dipole emission.

Note that now the gyration frequency does depend on the velocity (energy) of the
(relativistic) electrons, which in principle implies that because the electrons will
have a distribution in energies, the synchrotron emission is going to be continuum
emission. However, you can also see that the gyration frequency is even lower than
in the case of cyclotron emission, by a factor 1/γ. For the record, Lorentz factors of
up to ∼ 1011 have been measured, indicating that γ can be extremely large! Hence,
if the photon emission were to be at the gyration frequency, we would never be able
to see it, because of the plasme-frequency-shielding.

However, the gyration frequency is not the only frequency in this problem. Because
of the relativistic motion, the dipole emission from the electron, as seen from the
observer’s frame, is highly beamed (see Fig. 34), with an opening angle ∼ 1/γ (which
can thus be tiny). Consequently, the observer does not have a continuous view of the
electron, but only sees EM radiation when the beam sweeps over the line-of-sight.
The width of these ‘pulses’ are a factor 1/γ 3 shorter than the gyration period. The
corresponding frequency, called the critical frequency, is given by
3e
νcrit = γ 2 B⊥
4 π me c
which translates into
νcrit B⊥
= 4.2 γ 2
MHz Gauss

190
Figure 35: Specific intensity of synchrotron emission, including the effect of syn-
chrotron self absorption at low frequencies, where the optical depth exceeds unity.

So although the gyration frequency will be small, the critical frequency can be ex-
tremely large. This critical frequency corresponds to the shortest time period (the
pulse duration), and therefore represents the largest frequency, above which the emis-
sion is negligble. The longest time period, which is related to the gyration period,
determines the fundamental frequency

νf 2.8 ~
|B|
=
MHz γ sin2 φ Gauss

The emission spectrum due to synchrotron radiation will contain this fundamen-
tal frequency plus all its harmonics up to νcrit . Since these harmonics are ex-
tremely closely spaced (after all, the gyration frequency is extremely small), the
synchrotron spectrum for one value of γ looks essentially continuum. When taking
the γ-distribution into account (which is related to the energy distribution of the
relativistic electrons), the distribution becomes trully continuum, and the critical
and fundamental frequencies no longer can be discerned (because they depend on γ).

After integrating over the energy distribution of the relativistic electrons, which typ-
ically has a power-law distribution N(E) ∝ E −Γ one obtains the following emission

191
and absorption coefficients:
(Γ+1)/2
jν ∝ B⊥ ν −(Γ−1)/2
(Γ+2)/2
αν ∝ B⊥ ν −(Γ+4)/2

Note that αν describes synchrotron self-absorption. The resulting source func-


tion and optical depth are
jν −1/2
Sν = ∝ B⊥ ν 5/2
αν
(Γ+2)/2 −(Γ+4)/2
τν = αν l ∝ B⊥ ν l

Using that typically Γ > 0, we have that τν ∝ ν a with a < 0; synchrotron self-
absorption becomes more important at lower frequencies.

Application of the equation of radiative transfer, Iν = Sν (1 − e−τν ), yields



Sν ∝ ν 5/2 if τν ≫ 1
Iν =
jν l ∝ ν α if τν ≪ 1

where α ≡ − Γ−1 2
. Fig. 35 shown an illustration of a typical synchrotron spec-
trum: at low frequencies, where τν ≫ 1, we have that Iν ∝ ν 5/2 , which transits
to Iν ∝ ν −(Γ−1)/2 once the emitting medium becomes optically thin for synchroton
self-absorption. Note that there is no cut-off related to the critical frequency, since
νcrit = νcrit (E).

192
Supplemental Material

Appendices

Appendices A-E present background material on calculus relevant for this course.
The other Appendices present supplemental material that is NOT considered part of
this course’s curriculum. They are included to provide background information for
those readers that want to know a bit more.

Appendix A: Vector Calculus . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193


Appendix B: Conservative Vector Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 198
Appendix C: Integral Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199
Appendix D: Curvi-Linear Coordinate Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200
Appendix E: The Levi-Civita Symbol . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
Appendix F: The Viscous Stress Tensor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
Appendix G: The Chemical Potential . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214

193
Appendix A

Vector Calculus

~ = (a1 , a2 , a3 ) = a1 î + a2 ĵ + a3 k̂
Vector: A

p
~ =
Amplitude of vector: |A| a21 + a22 + a23

~ =1
Unit vector: |A|

Basis: In the above example, the unit vectors î, ĵ and k̂ form a vector basis.
~ B
Any 3 vectors A, ~ and C
~ can form a vector basis
~ B,
as long as det(A, ~ C)
~ 6= 0.

~ B)
~ a1 a2
Determinant: det(A, = = a1 b2 − a2 b1
b1 b2

a1 a2 a3
~ B,
~ C)
~ b2 b3 b3 b1 b1 b2
det(A, = b1 b2 b3 = a1 + a2 + a3
c2 c3 c3 c1 c1 c2
c1 c2 c3

Geometrically: ~ B)
det(A, ~ = ± area of parallelogram
~ B,
det(A, ~ C)
~ = ± volume of parallelepiped

Multiplication by scalar: αA~ = (αa1 , αa2 , αa3 )


~ = |α| |A|
|α A| ~

~+B
Summation of vectors: A ~ =B
~ +A
~ = (a1 + b1 , a2 + b2 , a3 + b3 )

194
P
Einstein Summation Convention: ai bi = i ai bi = a1 b1 + a2 b2 + a3 b3 = ~a · ~b
∂Ai /∂xi = ∂A1 /∂x1 + ∂A2 /∂x2 + ∂A3 /∂x3 = ∇ · A ~
Aii = A11 + A22 + A33 = Tr A ~ (trace of A)~

Dot product (aka scalar product): ~·B


A ~ = ai bi = |A|
~ |B|
~ cos θ
~·B
A ~ =B~ ·A ~
Useful for:
~ · B/(|
• computing angle between two vectors: cos θ = A ~ A| ~ |B|)
~

~ ·B
• check orthogonality: two vectors are orthogonal if A ~ =0

~ in direction of A,
• compute projection of B ~ which is given by A
~ · B/|
~ A|~

î ĵ k̂
~×B
Cross Product (aka vector product): A ~ = a1 a2 a3 = εijk ai bj êk
b1 b2 b3
~ ~ ~ |B|
|A × B| = |A| ~ sin θ = det(A,
~ B)
~

NOTE: εijk is called the Levi-Civita tensor, which is described in Appendix E,.

In addition to the dot product and cross product, there is a third vector product
that one occasionally encounters in dynamics;

Tensor product: ~⊗B


A ~ = AB (AB)ij = ai bj
~⊗B
A ~ 6= B
~ ⊗A
~

The tensor product AB is a tensor of rank two and is called a dyad. The sum of two
~ B,
or more dyads is called a dyadic. For example, let A, ~ C~ and D~ be four vectors,
from which we can form the dyads AB and CD. Their sum AB + CD is then a
dyadic. Note that in general a dyadic it not a dyad because it cannot be written
as a vector multiplied with a vector. Hence, dyadics differ from vectors in that the
sum of two vectors is a vector whereas the sum of two dyads is not necessarily a
dyad.

195
~·B
A ~ =B
~ ·A
~ ~×B
A ~ = −B
~ ×A
~

~ ·B
(αA) ~ = α(A
~ · B)
~ =A
~ · (αB)
~ ~ ×B
(αA) ~ = α(A
~ × B)
~ =A
~ × (αB)
~

~ · (B
A ~ + C)
~ =A
~·B
~ +A
~·C
~ ~ × (B
A ~ + C)
~ =A
~ ×B
~ +A
~×C
~

~·B
A ~ =0 → ~⊥B
A ~ ~ ×B
A ~ =0 → ~kB
A ~

~·A
A ~ = |A|
~2 ~ ×A
A ~=0

~ · (B
Triple Scalar Product: A ~ × C)
~ = det(A,
~ B,
~ C)
~ = εijk ai bj ck
~ · (B
A ~ × C)
~ = 0 → A, ~ B,
~ C
~ are coplanar
~ · (B
A ~ × C)
~ =B~ · (C
~ × A)
~ =C~ · (A~ × B)~

~ × (B
Triple Vector Product: A ~ × C)
~ = (A~ · C)
~ B~ − (A
~ · B)
~ C~
as is clear from above, A~ × (B
~ × C)
~ lies in plane of B
~ and C.
~

~ × B)
~ · (C
~ × D)
~ = (A
~ ~ ~ ~ ~ ~ ~ ~
Useful to remember: (A h · C) (B · D)i− (A ·hD) (B · C) i
~ × B)
(A ~ × (C~ × D)
~ = A ~ · (B
~ × D)
~ C ~− A ~ · (B
~ × C)
~ D ~

 
Gradient Operator: ∇ = ∇ ~ = ∂, ∂, ∂
∂x ∂y ∂z
This vector operator is sometimes called the nabla or del operator.

∂ 2 ∂ 2∂ 2
Laplacian operator: ∇2 = ∇ · ∇ = ∂x 2 + ∂y 2 + ∂z 2

This is a scalar operator.

∂f ∂f ∂f
Differential: f = f (x, y, z) → df = ∂x
dx + ∂y
dy + ∂z
dz

Chain Rule: If x = x(t), y = y(t) and z = z(t) then df dt


= ∂f dx
∂x dt
+ ∂f dy
∂y dt
+ ∂f dz
∂z dt
If x = x(s, t), y = y(s, t) and z = z(s, t) then ∂f
∂s
= ∂f ∂x
∂x ∂s
+ ∂f ∂y
∂y ∂s
+ ∂f ∂z
∂z ∂s

196
 
∂f ∂f ∂f
Gradient Vector: ∇f = gradf = , ,
∂x ∂y ∂z
the gradient vector at (x, y, z) is normal to the level surface
through the point (x, y, z).

Directional Derivative: The derivative of f = f (x, y, z) in direction of ~u is


Du f = ∇f · |~u~u| = |∇f | cos θ

Vector Field: F~ (~x) = (Fx , Fy , Fz ) = Fx î + Fy ĵ + Fz k̂


where Fx = Fx (x, y, z), Fy = Fy (x, y, z), and Fz = Fz (x, y, z).

Divergence of Vector Field: divF~ = ∇ · F~ = ∂F ∂x


x
+ ∂F
∂y
y
+ ∂F
∂z
z

A vector field for which ∇ · F~ = 0 is called solenoidal or divergence-free.

î ĵ k̂
Curl of Vector Field: curlF~ = ∇ × F~ = ∂/∂x ∂/∂y ∂/∂z
Fx Fy Fz
~
A vector field for which ∇ × F = 0 is called irrotational or curl-free.

Laplacian of Vector Field: ∇2 F~ = (∇ · ∇)F~ = ∇(∇ · F~ ) − ∇ × (∇ × F~ )


Note that ∇2 F~ =
6 ∇(∇ · F~ ): do not make this mistake.

~ x) and B(~
Let S(~x) and T (~x) be scalar fields, and let A(~ ~ x) be vector fields:

∇S = gradS = vector ∇2 S = ∇ · (∇S) = scalar

~ = divA
∇·A ~ = scalar ~ = (∇ · ∇) A
∇2 A ~ = vector

~ = curlA
∇×A ~ = vector

197
∇ × (∇S) = 0 curl grad S = 0

~ =0
∇ · (∇ × A) ~=0
div curl A

∇(ST ) = S ∇T + T ∇S

~ = S(∇ · A)
∇ · (S A) ~ +A
~ · ∇S

~ = (∇S) × A
∇ × (S A) ~ + S(∇ × A)
~

~ × B)
∇ · (A ~ =B
~ · (∇ × A)
~ −A
~ · (∇ × B)
~

~ × B)
∇ × (A ~ = A(∇
~ · B)
~ − B(∇
~ ~ + (B
· A) ~ · ∇)A
~ − (A
~ · ∇)B
~

~ · B)
∇(A ~ = (A
~ · ∇)B
~ + (B
~ · ∇)A
~ +A
~ × (∇ × B)
~ +B
~ × (∇ × A)
~

~ × (∇ × A)
A ~ = 1 ∇(A
~ · A)
~ − (A
~ · ∇)A
~
2

~ = ∇2 (∇ × A)
∇ × (∇2 A) ~

198
Appendix B

Conservative Vector Fields

Line Integral of a Conservative Vector Field: Consider a curve γ running from


location ~x0 to ~x1 . Let d~l be the directional element of length along γ (i.e., with
direction equal to that of the tangent vector to γ), then, for any scalar field Φ(~x),
Z ~
x1 Z ~
x1
∇Φ · d~l = dΦ = Φ(~x1 ) − Φ(~x0 )
~
x0 ~
x0

This implies that the line integral is independent of γ, and hence


I
∇Φ · d~l = 0
c

where c is a closed curve, and the integral is to be performed in the counter-clockwise


direction.

Conservative Vector Fields:


A conservative vector field F~ has the following properties:

• F~ (~x) is a gradient field, which means that there is a scalar field Φ(~x) so that
F~ = ∇Φ
H
• Path independence: c F~ · d~l = 0

• Irrotational = curl-free: ∇ × F~ = 0

199
Appendix C

Integral Theorems

Green’s Theorem: Consider a 2D vector field F~ = Fx î + Fy ĵ


I Z Z Z Z
F~ · d~l = ∇ × F~ · n̂ dA = |∇ × F~ | dA
A A
I Z Z
F~ · n̂ dl = ∇ · F~ dA
A

NOTE: in the first equation we have used that ∇ × F~ is always pointing in the
direction of the normal n̂.

Gauss’ Divergence Theorem: Consider a 3D vector field F~ = (Fx , Fy , Fz )


If S is a closed surface bounding a region D with normal pointing outwards, and F~
is a vector field defined and differentiable over all of D, then
Z Z Z Z Z
F~ · dS
~= ∇ · F~ dV
S D

Stokes’ Curl Theorem: Consider a 3D vector field F~ = (Fx , Fy , Fz )


If C is a closed curve, and S is any surface bounded by C, then
I Z Z
F~ · d~l = (∇ × F~ ) · n̂ dS
c S

NOTE: The curve of the line intergral must have positive orientation, meaning that
d~l points counterclockwise when the normal of the surface points towards the viewer.

200
Appendix D

Curvi-Linear Coordinate Systems


In astrophysics, one often works in curvi-linear, rather than Cartesian coordinate
systems. The two most often encountered examples are the cylindrical (R, φ, z)
and spherical (r, θ, φ) coordinate systems.

In this appendix we describe how to handle vector calculus in non-Cartesian coordi-


nate systems (Euclidean spaces only). After giving the ‘rules’ for arbitrary coordinate
systems, we apply them to cylindrical and spherical coordinate systems, respectively.

Vector Calculus in an Arbitrary Coordinate System:


Consider a vector ~x = (x, y, z) in Cartesian coordinates. This means that we can
write
~x = x ~ex + y ~ey + z ~ez
where ~ex , ~ey and ~ez are the unit directional vectors. Now consider the same vector
~x, but expressed in another general (arbitrary) coordinate system; ~x = (q1 , q2 , q3 ).
It is tempting, but terribly wrong, to write that

~x = q1 ~e1 + q2 ~e2 + q3 ~e3

where ~e1 , ~e2 and ~e3 are the unit directional vectors in the new (q1 , q2 , q3 )-coordinate
system. In what follows we show how to properly treat such generalized coordinate
systems.
In general, one expresses the distance between (q1 , q2 , q3 ) and (q1 + dq1 , q2 + dq2 , q3 +
dq3 ) in an arbitrary coordinate system as
p
ds = hij dqi dqj

Here hij is called the metric tensor. In what follows, we will only consider orthog-
onal coordinate systems for which√hij = 0 if i 6= j, so that ds2 = h2i dqi2 (Einstein
summation convention) with hi = hii .
An example of an orthogonal coordinate system are the Cartesian coordinates, for
which hij = δij . After all, the distance between two points separated by the infinites-
imal displacement vector d~x = (dx, dy, dz) is ds2 = |d~x|2 = dx2 + dy 2 + dz 2 .

201
The coordinates (x, y, z) and (q1 , q2 , q3 ) are related to each other via the transfor-
mation relations
x = x(q1 , q2 , q3 )
y = y(q1 , q2 , q3 )
z = z(q1 , q2 , q3 )

and the corresponding inverse relations


q1 = q1 (x, y, z)
q2 = q2 (x, y, z)
q3 = q3 (x, y, z)

Hence, we have that the differential vector is:


∂~x ∂~x ∂~x
d~x = dq1 + dq2 + dq3
∂q1 ∂q2 ∂q3
where
∂~x ∂
= (x, y, z)
∂qi ∂qi
The unit directional vectors are:
∂~x/∂qi
~ei =
|∂~x/∂qi |
which allows us to rewrite the expression for the differential vector as
∂~x ∂~x ∂~x
d~x = dq1 ~e1 + dq2 ~e2 + dq3 ~e3
∂q1 ∂q2 ∂q3
and thus
2
∂~x
|d~x|2 = dqi2
∂qi
(Einstein summation convention). Using the definition of the metric, according to
which |d~x|2 = h2i dqi2 we thus infer that

∂~x
hi =
∂qi

202
Using this expression for the metric allows us to write the unit directional vectors
as
1 ∂~x
~ei =
hi ∂qi
and the differential vector in the compact form as

d~x = hi dqi ~ei

From the latter we also have that the infinitesimal volume element for a general
coordinate system is given by

d3~x = |h1 h2 h3 | dq1 dq2 dq3

Note that the absolute values are needed to assure that d3~x is positive.

203
~ In the Cartesian basis C = {~ex , ~ey , ~ez } we have that
Now consider a vector A.
~ C = Ax ~ex + Ay ~ey + Az ~ez
[A]
In the basis B = {~e1 , ~e2 , ~e3 }, corresponding to our generalized coordinate system,
we instead have that
~ B = A1 ~e1 + A2 ~e2 + A3 ~e3
[A]
We can rewrite the above as
       
e11 e21 e31 A1 e11 + A2 e21 + A3 e31
~ B = A1  e12  + A2  e22  + A3  e32  =  A2 e12 + A2 e22 + A3 e32 
[A]
e13 e23 e33 A3 e13 + A2 e23 + A3 e33
and thus     
e11 e21 e31 A1 A1
~ B =  e12 e22 e32   A2  ≡ T  A2 
[A]
e13 e23 e33 A3 A3
Using similar logic, one can write
       
ex1 ey1 ez1 Ax 1 0 0 Ax Ax
~ C =  ex2 ey2 ez2   Ay  =  0 1 0   Ay  = I  Ay 
[A]
ex3 ey3 ez3 Az 0 0 1 Az Az
~ is the same object independent of its basis we have that
and since A
   
Ax A1
I  Ay  = T  A2 
Az A3
~ B and [A]
and thus, we see that the relation between [A] ~ C is given by
~ C = T [A]
[A] ~ B, ~ B = T−1 [A]
[A] ~C

For this reason, T is called the transformation of basis matrix. Note that the
columns of T are the unit-direction vectors ~ei , i.e., Tij = eij . Since these are or-
thogonal to each other, the matric T is said to be orthogonal, which implies that
T−1 = T T (the inverse is equal to the transpose), and det(T ) = ±1.
Now we are finally ready to determine how to write our position vector ~x in the new
basis B of our generalized coordinate system. Let’s write ~x = ai ~ei , i.e.
 
a1
[~x]B =  a2 
a3

204
We started this appendix by pointing out that it is tempting, but
p wrong, to set ai = qi
(as for the Cartesian basis). To see this, recall that |~x| = (a1 )2 + (a2 )2 + (a3 )2 ,
from which it is immediately clear that each ai needs to have the dimension of length.
Hence, when qi is an angle, clearly ai 6= qi . To compute the actual ai you need to
use the transformation of basis matrix as follows:
    
e11 e12 e13 x e11 x + e12 y + e13 z
[~x]B = T−1 [~x]C =  e21 e22 e23   y  =  e21 x + e22 y + e23 z 
e31 e32 e33 z e31 x + e32 y + e33 z

Hence, using our expression for the unit direction vectors, we see that
   
1 ∂xj 1 ∂~x
ai = xj = · ~x
hi ∂qi hi ∂qi

Hence, the position vector in the generalized basis B is given by


X 1  ∂~x 
[~x]B = · ~x ~ei
i
hi ∂qi

and by operating d/dt on [~x]B we find that the corresponding velocity vector in the
B basis is given by X
[~v ]B = hi q̇i ~ei
i

with q̇i = dqi /dt. Note that the latter can also be inferred more directly by simply
dividing the expression for the differential vector (d~x = hi qi ~ei ) by dt.

205
Next we write out the gradient, the divergence, the curl and the Laplacian for our
generalized coordinate system:

The gradient:
1 ∂ψ
∇ψ = ~ei
hi ∂qi

The divergence:
 
~= 1 ∂ ∂ ∂
∇·A (h2 h3 A1 ) + (h3 h1 A2 ) + (h1 h2 A3 )
h1 h2 h3 ∂q1 ∂q2 ∂q3

The curl (only one component shown):


 
~ 1 ∂ ∂
(∇ × A)3 = (h2 A2 ) − (h1 A1 )
h1 h2 ∂q1 ∂q2

The Laplacian:
      
2 1 ∂ h2 h3 ∂ψ ∂ h3 h1 ∂ψ ∂ h1 h2 ∂ψ
∇ ψ= + +
h1 h2 h3 ∂q1 h1 ∂q1 ∂q2 h2 ∂q2 ∂q3 h3 ∂q3

The Convective operator:


  
~ · ∇) B
~ = Ai ∂Bj B i ∂hj ∂hi
(A + Aj − Ai ~ej
hi ∂qi hi hj ∂qi ∂qj

206
Vector Calculus in Cylindrical Coordinates:

For cylindrical coordinates (R, φ, z) we have that

x = R cos φ y = R sin φ z=z

The scale factors of the metric therefore are:

hR = 1 hφ = R hz = 1

and the position vector is ~x = R~eR + z~ez .

~ = AR~eR + Aφ~eφ + Az ~ez an arbitrary vector, then


Let A

AR = Ax cos φ − Ay sin φ
Aφ = −Ax sin φ + Ay cos φ
Az = Az

In cylindrical coordinates the velocity vector becomes:

~v = Ṙ ~eR + R ~e˙ R + ż ~ez


= Ṙ ~eR + R φ̇ ~eφ + ż ~ez

The Gradient:
~ = 1 ∂ (RAR ) + 1 ∂Aφ + ∂Az
∇·A
R ∂R R ∂φ ∂z

The Convective Operator:


 
~ · ∇) B
~ = ∂B R Aφ ∂BR ∂B R Aφ Bφ
(A AR + + Az − ~eR
∂R R ∂φ ∂z R
 
∂Bφ Aφ ∂Bφ ∂Bφ Aφ BR
+ AR + + Az + ~eφ
∂R R ∂φ ∂z R
 
∂Bz Aφ ∂Bz ∂Bz
+ AR + + Az ~ez
∂R R ∂φ ∂z

207
The Laplacian:
 
2 1 ∂ ∂ψ 1 ∂2ψ ∂2ψ
scalar : ∇ψ = R + + 2
R ∂R ∂R R2 ∂φ2 ∂z

 
2~ 2 FR 2 ∂Fθ
vector : ∇ F = ∇ FR − 2 − 2 ~eR
R R ∂θ
 
2 2 ∂FR Fθ
+ ∇ Fθ + 2 − 2 ~eθ
R ∂θ R
2

+ ∇ Fz ~ez

208
Vector Calculus in Spherical Coordinates:

For spherical coordinates (r, θ, φ) we have that

x = r sin θ cos φ y = r sin θ sin φ z = r cos θ

The scale factors of the metric therefore are:

hr = 1 hθ = r hφ = r sin θ

and the position vector is ~x = r~er .

~ = Ar~er + Aθ~eθ + Aφ~eφ an arbitrary vector, then


Let A

Ar = Ax sin θ cos φ + Ay sin θ sin φ + Az cos θ


Aθ = Ax cos θ cos φ + Ay cos θ sin φ − Az sin θ
Aφ = −Ax sin φ + Ay cos φ

In spherical coordinates the velocity vector becomes:

~v = ṙ ~er + r ~e˙ r
= ṙ ~er + r θ̇ ~eθ + r sin θ φ̇ ~eφ

The Gradient:

~ = 1 ∂ (r 2 Ar ) + 1
∇·A

(sin θAθ ) +
1 ∂Aφ
2
r ∂r r sin θ ∂θ r sin θ ∂φ
The Convective Operator:
 
~ ~ ∂Br Aθ ∂Br Aφ ∂Br Aθ Bθ + Aφ Bφ
(A · ∇) B = Ar + + − ~er
∂r r ∂θ r sin θ ∂φ r
 
∂Bθ Aθ ∂Bθ Aφ ∂Bθ Aθ Br Aφ Bφ cotθ
+ Ar + + + − ~eθ
∂r r ∂θ r sin θ ∂φ r r
 
∂Bφ Aθ ∂Bφ Aφ ∂Bφ Aφ Br Aφ Bθ cotθ
+ Ar + + + + ~eφ
∂r r ∂θ r sin θ ∂φ r r

209
The Laplacian:
   
2 1 ∂ 2 ∂ψ 1 ∂ ∂φ 1 ∂2ψ
scalar : ∇ψ = 2 r + 2 sin θ + 2 2
r ∂r ∂r r sin θ ∂θ ∂θ r sin θ ∂ψ 2

 
2~ 2 2Fr 2 ∂(Fθ sin θ) 2 ∂Fφ
vector : ∇F = ∇ Fr − 2 − 2 − 2 ~er
r r sin θ ∂θ r sin θ ∂φ
 
2 2 ∂Fr Fθ 2 cos θ ∂Fφ
+ ∇ Fθ + 2 − 2 − ~eθ
r ∂θ r sin θ r 2 sin2 θ ∂φ
 
2 2 ∂Fr 2 cos θ ∂Fθ Fφ
+ ∇ Fφ + 2 + 2 2 − 2 2 ~eφ
r sin θ ∂φ r sin θ ∂φ r sin θ

210
Appendix E

The Levi-Civita Symbol

The Levi-Civita symbol, also known as the permutation symbol or the anti-
symmetric symbol,is a collection of numbers, defined from the sign of a permu-
tation of the natural numbers 1, 2, 3, ..., n. It is often encountered in linear algebra,
vector and tensor calculus, and differential geometry.

The n-dimensional Levi-Civita symbol is indicated by εi1 i2 ...in , where each index
i1 , i2 , ..., in takes values 1, 2, ..., n, and has the defining property that the symbol is
total antisymmetric in all its indices: when any two indices are interchanged, the
symbol is negated:
ε...ip...iq ... = −ε...iq ...ip ...
If any two indices are equal, the symbol is zero, and when all indices are unequal,
we have that
εi1 i2 ...in = (−1)p ε1,2,...n
where p is called the parity of the permutation. It is the number of pairwise inter-
changes necessary to unscramble i1 , i2 , ..., in into the order 1, 2, ..., n. A permutation
is said to be even (odd) if its parity is an even (odd) number.

Example: what is the parity of {3, 4, 5, 2, 1}?


{1, 2, 3, 4, 5}
{3, 2, 1, 4, 5}
{3, 4, 1, 2, 5}
{3, 4, 5, 2, 1}
Answer: p = 3, since three pairwise interchanges are required.

In three dimensions the Levi-Civita symbol is defined by



 +1 if (i, j, k) is (1,2,3), (2,3,1), or (3,1,2)
εijk = −1 if (i, j, k) is (3,2,1), (1,3,2), or (2,1,3)

0 if i = j, or j = k, or k = i

211
Appendix F

The Viscous Stress Tensor


As discussed in Chapter 4, the deviatoric stress tensor, τij , is only non-zero in the
presence of shear in the fluid flow. This suggests that
∂uk
τij = Tijkl
∂xl
where Tijkl is a proportionality tensor of rank four. In what follows we derive an
expression for Tijkl . We start by noting that since σij is symmetric, we also have
that τij will be symmetric. Hence, we expect that the above dependence can only
involve the symmetric component of the deformation tensor, Tkl = ∂uk /∂xl . Hence,
it is useful to split the deformation tensor in its symmetric and anti-symmetric
components:
∂ui
= eij + ξij
∂xj
where

 
1 ∂ui ∂uj
eij = +
2 ∂xj ∂xi
 
1 ∂ui ∂uj
ξij = −
2 ∂xj ∂xi

The symmetric part of the deformation tensor, eij , is called the rate of strain
tensor, while the anti-symmetric part, ξij , expresses the vorticity w ~ ≡ ∇ × ~u in
1
the velocity field, i.e., ξij = − 2 εijk wk . Note that one can always find a coordinate
system for which eij is diagonal. The axes of that coordinate frame indicate the
eigendirections of the strain (compression or stretching) on the fluid element.

In terms of the relation between the viscous stress tensor, τij , and the deformation
tensor, Tkl , there are a number of properties that are important.

212
• Locality: the τij − Tkl -relation is said to be local if the stress tensor is only
a function of the deformation tensor and thermodynamic state functions like
temperature.

• Homogeneity: the τij − Tkl -relation is said to be homogeneous if it is ev-


erywhere the same. The viscous stress tensor may depend on location ~x only
insofar as Tij or the thermodynamic state functions depend on ~x. This distin-
guishes a fluid from a solid, in which the stress tensor depends on the stress
itself.

• Isotropy: the τij − Tkl -relation is said to be isotropic if it has no preferred


direction.

• Linearity: the τij − Tkl -relation is said to be linear if the relation between
the stress and rate-of-strain is linear. This is equivalent to saying that τij does
not depend on ∇2~u or higher-order derivatives.

A fluid that is local, homogeneous and isotropic is called a Stokesian fluid. A


Stokesian fluid that is linear is called a Newtonian fluid. Experiments have shown
that most (astrophysical) fluids are Newtonian to good approximation. Hence, in
what follows we will assume that our fluids are Newtonian, unless specifically stated
otherwise. For a Newtonian fluid, it can be shown (using linear algebra) that the
most general form of our proportionality tensor is given by

Tijkl = λδij δkl + µ (δik δjl + δil δjk )


Hence, for a Newtonian fluid the viscous stress tensor is

τij = 2µeij + λekk δij


where µ is the coefficient of shear viscosity, λ is a scalar, δij is the Kronecker
delta function, and ekk = Tr(e) = ∂uk /∂xk = ∇ · ~u (summation convention).

Note that (in a Newtonian fluid) the viscous stress tensor depends only on the sym-
metric component of the deformation tensor (the rate-of-strain tensor eij ), but not
on the antisymmetric component which describes vorticity. You can understand
the fact that viscosity and vorticity are unrelated by considering a fluid disk in solid
body rotation (i.e., ∇ · ~u = 0 and ∇ × ~u = w ~ 6= 0). In such a fluid there is no
”slippage”, hence no shear, and therefore no manifestation of viscosity.

213
Thus far we have derived that the stress tensor, σij , which in principle has 6 un-
knowns, can be reduced to a function of three unknowns only (P , µ, λ) as long as
the fluid is Newtonian. Note that these three scalars, in general, are functions of
temperature and density. We now focus on these three scalars in more detail, starting
with the pressure P . To be exact, P is the thermodynamic equilibrium pres-
sure, and is normally computed thermodynamically from some equation of state,
P = P (ρ, T ). It is related to the translational kinetic energy of the particles when
the fluid, in equilibrium, has reached equipartition of energy among all its degrees
of freedom, including (in the case of molecules) rotational and vibrations degrees of
freedom.

In addition to the thermodynamic equilibrium pressure, P , we can also define a


mechanical pressure, Pm , which is purely related to the translational motion of
the particles, independent of whether the system has reached full equipartition of
energy. The mechanical pressure is simply the average normal stress and therefore
follows from the stress tensor according to
1 1
Pm = − Tr(σij ) = − (σ11 + σ22 + σ33 )
3 3
Using that

σij = −P δij + 2 µ eij + λ ekk δij


we thus obtain the following relation between the two pressures:

Pm = P − η ∇ · ~u
where
2 P − Pm
η = µ+λ=
3 ∇ · ~u
is the coefficient of bulk viscosity. We can now write the stress tensor as

 
∂ui ∂uj 2 ∂uk ∂uk
σij = −P δij + µ + − δij + η δij
∂xj ∂xi 3 ∂xk ∂xk

This is the full expression for the stress tensor in terms of the coefficients of shear
viscosity, µ, and bulk viscosity, η.

214
Appendix G

The Chemical Potential

Consider a system which can exchange energy and particles with a reservoir, and
the volume of which can change. There are three ways for this system to increase
its internal energy; heating, changing the system’s volume (i.e., doing work on the
system), or adding particles. Hence,

dU = T dS − P dV + µ dN

Note that this is the first law of thermodynamics, but now with the added possibility
of changing the number of particles of the system. The scalar quantity µ is called
the chemical potential, and is defined by
 
∂U
µ=
∂N S,V

This is not to be confused with the µ used to denote the mean weight per particle,
which ALWAYS appears in combination with the proton mass, mp . As is evident
from the above expression, the chemical potential quantifies how the internal energy
of the system changes if particles are added or removed, while keeping the entropy
and volume of the system fixed. The chemical potential appears in the Fermi-Dirac
distribution describing the momentum distribution of a gas of fermions or bosons.

Consider an ideal gas, of volume V , entropy S and with internal energy U. Now
imagine adding a particle of zero energy (ǫ = 0), while keeping the volume fixed.
Since ǫ = 0, we also have that dU = 0. But what about the entropy? Well, we
have increased the number of ways in which we can redistribute the energy U (a
macrostate quantity) over the different particles (different microstates). Hence, by
adding this particle we have increased the system’s entropy. If we want to add a
particle while keeping S fixed, we need to decrease U to offset the increase in the
number of ‘degrees of freedom’ over which to distribute this energy. Hence, keeping
S (and V ) fixed, requires that the particle has negative energy, and we thus see that
µ < 0.

215
For a fully degenerate Fermi gas, we have that T = 0, and thus S = 0 (i.e., there is
only one micro-state associated with this macrostate, and that is the fully degenerate
one). If we now add a particle, and demand that we keep S = 0, then that particle
must have the Fermi energy (see Chapter 6); ǫ = Ef . Hence, for a fully degenerate
gas, µ = Ef .

Finally, consider a photon gas in thermal equilibrium inside a container. Contrary to


an ideal gas, in a photon gas the number of particles (photons) cannot be arbitrary.
The number of photons at given temperature, T , and thus at given U, is given by
the Planck distribution and is constantly adjusted (through absorption and emission
against the wall of the container) so that the photon gas remains in thermal equilib-
rium. In other words, Nγ is not a degree of freedom for the system, but it set by the
volume and the temperature of the gas. Since we can’t change N while maintaining
S (or T ) and V , we have that µ = 0 for photons.

To end this discussion of the chemical potential, we address the origin of its name,
which may, at first, seem weird. Let’s start with the ‘potential’ part. The origin of
this name is clear from the following. According to its definition (see above), the
chemical potential is the ‘internal energy’ per unit amount (moles). Now consider
the following correspondences:

Gravitational potential is the gravitational energy per unit mass:


G m1 m2 Gm ∂W
W = ⇒ φ= ⇒ φ=
r r ∂m

Similarly, electrical potential is the electrical energy per unit charge


1 q1 q2 1 q ∂V
V = ⇒ φ= ⇒ φ=
4πε0 r 4πε0 r ∂q

These examples make it clear why µ is considered a ‘potential’. Finally, the word
chemical arises from the fact that the µ plays an important role in chemistry (i.e.,
when considering systems in which chemical reactions take place, which change the
particles). In this respect, it is important to be aware of the fact that µ is an
additive quantity that is conserved in a chemical reaction. Hence, for a chemical

216
reaction i + j → k + l one has that µi + µj = µk + µl . As an example, consider the
annihilation of an electron and a positron into two photons. Using that µ = 0 for
photons, we see that the chemical potential of elementary particles (i.e., electrons)
must be opposite to that of their anti-particles (i.e., positrons).

Because of the additive nature of the chemical potential, we also have that the
above equation for dU changes slightly whenever the gas consists of different particle
species; it becomes X
dU = T dS − P dV + µi dNi
i

where the summation is over all species i. If the gas consists of equal numbers
of elementary particles and anti-particles, then the total chemical potential of the
system will be P equal to zero. In fact, in many treatments of fluid dynamics it may be
assumed that i µi dNi = 0; in particular when the relevant reactions are ‘frozen’
(i.e., occur on a timescales τreact that are much longer than the dynamical timescales
τdyn of interest), so that dNi = 0, or if the reactions go so fast (τreact ≪ τdyn )
that each
P reaction and its inverse are in local thermodynamic equilibrium, in which
case i µi dNi = 0 for those species involved in the reaction. Only in the rare,
intermediate case when τreact ∼ τdyn is it important to keep track of the relative
abundances of the various chemical and/or nuclear species.

217

You might also like