0% found this document useful (0 votes)
17 views99 pages

Physics IV: Light and Quantum Mechanics

The document consists of lecture notes on the Fundamentals of Physics IV, covering topics such as light, particles, atoms, and quantum mechanics. It discusses wave properties of light, including interference and diffraction, and details experiments like Young's double-slit experiment. Additionally, it explores the mathematical foundations of quantum mechanics, including state vectors and operators.

Uploaded by

Thảo Ly
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views99 pages

Physics IV: Light and Quantum Mechanics

The document consists of lecture notes on the Fundamentals of Physics IV, covering topics such as light, particles, atoms, and quantum mechanics. It discusses wave properties of light, including interference and diffraction, and details experiments like Young's double-slit experiment. Additionally, it explores the mathematical foundations of quantum mechanics, including state vectors and operators.

Uploaded by

Thảo Ly
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Fundamentals of Physics IV

Lecture notes

written by
István Nándori and Zoltán Trócsányi

Debrecen, 2019
Contents

1 Light 3
1.1 Wave properties of light . . . . . . . . . . . . . . . . . 3
1.1.1 Young’s double-slit experiment . . . . . . . . . 4
1.1.2 Intensity in the interference pattern produced
in the double-slit experiment . . . . . . . . . . 6
1.1.3 Thin-film interference . . . . . . . . . . . . . . 7
1.1.4 Single-slit diffraction . . . . . . . . . . . . . . . 9
1.1.5 Intensity in single-slit diffraction . . . . . . . . 11
1.1.6 Diffraction (or precisely interference) gratings . 14
1.2 Blackbody radiation . . . . . . . . . . . . . . . . . . . 17
1.3 Particle properties of light . . . . . . . . . . . . . . . . 23
1.3.1 Photoelectric effect . . . . . . . . . . . . . . . . 23
1.3.2 Compton scattering . . . . . . . . . . . . . . . 26

2 Particles 31
2.1 Wave properties of matter . . . . . . . . . . . . . . . . 31
2.1.1 X-ray diffraction . . . . . . . . . . . . . . . . . 32
2.1.2 Diffraction of electron beam on crystals . . . . 34
2.1.3 Double-slit experiment with electrons . . . . . 34

3 Atoms 41
3.1 Rutherford’s experiment . . . . . . . . . . . . . . . . . 41
3.1.1 Setup and result . . . . . . . . . . . . . . . . . 41
3.1.2 Algebraic properties of a hyperbola . . . . . . . 43
3.1.3 Motion in an 1/r repulsive potential . . . . . . 45
3.1.4 Differential cross section of scattering on a point-
like scattering centre . . . . . . . . . . . . . . . 48
1
2 CONTENTS

3.1.5 Differential cross section of scattering on an ex-


tended scattering centre . . . . . . . . . . . . . 49
3.1.6 Significance of the discovery of the nucleus . . 51
3.2 Atomic spectra . . . . . . . . . . . . . . . . . . . . . . 52
3.2.1 Quantized energy levels . . . . . . . . . . . . . 57
3.2.2 Quantized angular momentum . . . . . . . . . 60
3.3 Splittings of spectral lines . . . . . . . . . . . . . . . . 61
3.3.1 Fine structure . . . . . . . . . . . . . . . . . . . 61
3.3.2 Zeeman effect . . . . . . . . . . . . . . . . . . . 62
3.3.3 Stern-Gerlach experiment . . . . . . . . . . . . 64
3.3.4 Stark effect . . . . . . . . . . . . . . . . . . . . 67
3.4 X-ray spectrum . . . . . . . . . . . . . . . . . . . . . . 68
3.4.1 Interpretation of Moseley’s plot . . . . . . . . . 70
3.5 Light amplification by the stimulated emission of radi-
ation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
3.6 Atomic quantum numbers . . . . . . . . . . . . . . . . 74
3.6.1 Magnitude of the angular momentum . . . . . 75
3.6.2 Direction of the angular momentum . . . . . . 75
3.6.3 Intrinsic angular momentum . . . . . . . . . . 77
3.7 Atoms . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
3.8 Periodic table of elements . . . . . . . . . . . . . . . . 80

4 Basics of Quantum Mechanics 83


4.1 Predictions in quantum mechanics . . . . . . . . . . . 83
4.2 Mathematical model of quantum mechanics . . . . . . 85
4.2.1 Physical states as vectors . . . . . . . . . . . . 85
4.2.2 Representations of state vectors . . . . . . . . . 88
4.2.3 Physical quantities as operators . . . . . . . . . 89
4.2.4 Representation of operators . . . . . . . . . . . 91
4.3 Heisenberg’s uncertainty principle . . . . . . . . . . . . 96
4.2.1 Physical states as vectors:
In quantum mechanics, physical states are represented as vectors in a Hilbert space. This section discusses the properties and characteristics of these
state vectors. It covers topics such as normalization, superposition, and the concept of inner product between state vectors.

4.2.2 Representations of state vectors:


To perform calculations and make predictions in quantum mechanics, it is often necessary to represent state vectors in a specific basis. This section
explores different representations of state vectors, such as the position representation and the momentum representation. It explains how to express
state vectors using basis functions and discusses the concept of completeness.

4.2.3 Physical quantities as operators:


In quantum mechanics, physical quantities, such as position, momentum, and energy, are represented by mathematical operators. This section
introduces the concept of operators and their role in quantum mechanics. It explains how operators act on state vectors to obtain measurable quantities
and covers important properties of operators, including linearity, Hermiticity, and eigenvalues.

4.2.4 Representation of operators:


Similar to state vectors, operators can be represented in different bases. This section focuses on the representation of operators using matrices. It
explains how the matrix elements of operators correspond to expectation values and probabilities of measurement outcomes. It also discusses important
operators in quantum mechanics, such as the position operator, momentum operator, and Hamiltonian operator.
Chapter 1

Light

1.1 Wave properties of light


Light plays a central role in life and also in physics. It is the prime
means of collecting information about our environment, so under-
standing its nature and properties is an important task in physics.
The first pieces of observations about light were the following:
1. Light travels at constant speed in empty speed, independently of
the speed of its source, which lead us to the principle of special
relativity.
2. In medium the speed of light is always smaller than it is in empty
space. The fraction of decrease depends on the medium and is
called the refraction index. Furthermore, in a(n optically) ho-
mogeneous medium light travels in straight lines that we called
light rays. These observations lead us to develop the laws of
geometrical optics.
3. Near the edges of barriers light does not travel in straight lines,
we observe diffraction similarly as in the case of waves on the
surface of water, for instance in a ripple tank. This was the first
hint that light shows properties of waves.
In addition to the observation of diffraction, we can test further
whether light indeed behaves as waves by putting two narrow slits
in the way of a laser beam. We can indeed observe the interference
pattern observed in the case of the ripple tank, which confirms that
3
4 CHAPTER 1. LIGHT

light is a wave. Note that we use laser beam in such an experiment.


In order to observe an interference pattern, the waves must arrive
at a fixed position in space with constant phase difference in time.
If we add such waves linearly, as required by the principle of linear
superposition, they will provide a wave intensity that is independent
of time at the given point in space. The phase difference may vary
at different locations, leading to different resulting intensity. Such
waves are said to be coherent. Coherent light can be produced if we
let the light of an ordinary light source, such as a light bulb, through
a tiny slit, and subsequently through two, or more small slits. The
only problem with such sources of coherent light is the small intensity
which makes observation difficult. With the invention of the laser,
much higher intensity coherent light source has become available.
There are two characteristics of waves: diffraction and interference.
Both can be understood on the basis of Huygens-Fresnel principle:
[Link]
why-does-light-bend-around-corners/v/huygen-s-principle-of-secondary-waves
All points on a wavefront can be considered as point sources
for the production of spherical secondary wavelets. After
a time t the new wave pattern is an interference of these
secondary wavelets.

In order to make a complete description of general cases, one has to


clarify the nature of light waves. In our course on electromagnetism,
we argued that light was electromagnetic wave of very specific wave-
length, falling into the range [400,800] nm. In principle, this makes
possible the description of any light, but that requires high level math-
ematics. In order to understand the physics of light, we can avoid
complicated mathematics if we concentrate on simple cases.

1.1.1 Young’s double-slit experiment


In 1801 Thomas Young was the first to perform the experiment de-
scribed above, applying two subsequent barriers with one and two
small slits on them. According to Huygens’ principle, the first slit
provides spherical waves that appear as almost plane waves when they
arrive on the second barrier containing the two slits. As we saw in
class, the same can be achieved using a laser beam. Thus the experi-
mental setup is essentially plane waves arriving on a barrier with two
slits on it, and we observe the resulting wave pattern on a screen at a
large distance as compared to the size and distance of the slits. We can
consider the wave pattern on a particular position P , characterized
1
The light rays are in phase at the slits because they derive 600 # 10"9 m. Then Eq. 35-19 tells us that, to shift the lower
from the same wave,difference
The phase but their relative
between twophase canchange
waves can shift on thewaves m
if the travel
= 1paths of fringe up to the center of the interference pat-
bright
way to the different
screen due to (1) a difference in the length of the
lengths.
tern, the plastic must have the thickness
paths they follow and (2) a difference in the number of their
internal l(N2 " N1) (600 # 10"9 m)(1)
Thewavelengths ln indifference
change in phase the materials through
is due which
to the path theydifference !LLin!the
length !
pass. The first
paths condition
taken applies
by the waves. to anytwo
Consider off-center
waves initially and in phase, traveling n2 " n1
point,exactly 1.50 " 1.00
the second
along condition
paths with applies when difference
a path length the plastic!L,
covers
and one
then of passing through some
! 1.2 # 10"6 m. (Answer)
CHAPTER 1. LIGHT
the slits.
common point. When !L is zero or an integer number of wavelengths, the waves 5
arrive at the common point exactly in phase and they interfere fully con-
Path length difference:
structively there. If thatFigure 35-11a
is true for theshows
wavesrays r1 and
of rays r2 r2 in Fig. 35-10, then
r1 and
along which waves from the two slits travel to reach the
lower m ! 1 bright fringe. D Those waves start in phase at the The difference in indexes
slits but arrive at the fringe with a phase difference of causes a phase shift
exactly 1 wavelength. To remind ourselves of this main r 2 char- The ∆L shifts
between the rays, moving
acteristic of the fringe, let us call it the 1l fringe. The one-
Incident P
one wave from the 1l fringe upward.
wave r
wavelength phase differencer is due to the one-wavelength
2
θ the other, which
1 S2
path length difference between theyrays reaching the fringe; determines the
m=1
that is, thereS 2 is exactly one d more wavelength d
θ along rayr 1 r2 interference.
r2
than along r1S.1 b θ Figure = 0 (a) Waves from slits S1 and S21l fringe
m 35-10
r2
Figure 35-11b shows the 1l fringeS 1shifted θ b up to the (which extend into and out of the r 1 page)
Path length difference ∆L combine at P, an arbitrary point on screen
center of the pattern with the plastic strip (b )
over the top slit r1 m=1
C at1ldistance
fringe y from the central axis. The
(we still do not know whether the plastic should be there angle u serves as a convenient locator for P.
or over the bottom slit). The figure also shows the new ori- (b) For D " d, we can approximate rays r1
entations of rays r1 and r2 to reach that fringe. There still and r2 as being parallel, at angle u to the
must be(a ) one Bmore wavelength along C r2 than along r1 (be- (a) central axis. (b)
cause they still produce the 1l fringe), but now the path Figure 35-11 (a) Arrangement for two-slit interference (not to scale).
length difference between those rays is zero, as we can tell The locations of three bright fringes (or maxima) are indicated.
from the geometry of Fig. 35-11b. However, r2 now passes (b) A strip of plastic covers the top slit. We want the 1l fringe to be
through the plastic. at the center of the pattern.
Figure 1.1:
Additional examples, video, and practice available at WileyPLUS

by the angle θ on the screen (see Fig. 1.1) as the linear superposi-
tion of two waves starting at the slits of the second barrier with the
same phase, but travelling different distances, therefore, arriving with
a phase difference. The latter is due to the different distance that
can be computed by simple geometry. We assume that the distance
between the two slits d is much smaller than the distance from the
barrier to the screen D, d << D and the position of observation on
the screen P is at a distance r << D from the closest position on the
screen O to the slits. We denote the angle arctan r/D by θ. The path
difference from the two slits to P is approximately d sin θ. If this path
difference is equal to integer number n times the wavelength, then the
resulting phase difference is zero and the two waves add constructively.
Thus at distances r, satisfying the condition

d sin θ = mλ , (1.1.1)

there will be lines of maximum intensity. If however, the path differ-


ence is equal to half integer number times the wavelength, then the
resulting phase difference is π and the two waves add destructively.
Thus at distances r, satisfying the condition

d sin θ = (m + 1/2)λ , (1.1.2)

there will be lines of total cancellation between the two light waves,
(minimum intensity).
6 CHAPTER 1. LIGHT

1.1.2 Intensity in the interference pattern produced


in the double-slit experiment
Eqs. (1.1.1) and (1.1.2) tell us the positions of the maxima and minima
as a function of the angle θ in the double slit experiment. In the class
we also saw that the brightness of the fringes depends on the angle.
Let us try to the derive an expression for the intensity I of the fringes
as a function of θ.
We mentioned that Young created coherent waves using a single
narrow slit as a source. Thus, ideally we can assume that the waves
leaving the two slits are in phase. When they arrive at point P on
the screen they are not in phase any longer. The electric field in
the light originating from slit i is Ei = E0 sin(ωt + φi ) where ω is the
angular frequency of the wave and φi is the phase constant of the wave
from slit i, while the amplitudes of the fields are equal. The phase
difference is φ = φ2 − φ1 . The phase difference does not change with
time, which means that the waves are coherent. It is proportional to
the path difference ∆` according to the relation

∆`
φ = 2π (1.1.3)
λ

where λ is the wavelength of light. As discussed above, the path


difference is approximately d sin θ, so

d
φ ≈ 2π sin θ . (1.1.4)
λ

We can represent the field Ei as the projection onto the vertical


axis of a vector of length E0 rotating in the plane orthogonal to the
direction of propagation of light with angular speed ω. The scalar
product of these vectors is a constant E02 cos φ, i.e. they rotate with
a fixed angle, the phase difference between them. These vectors are
called phasors. The net electric field will be the projection of the sum
of these vectors onto the vertical axis. This vector also rotates with
the same angular speed. As the vectors Ei have the same length, upon
addition they form an equilateral triangle, with base being the sum
of the vectors. Thus the phase of the sum is ψ = φ/2 and its length
is E = 2E0 cos ψ (see Fig. 1.2), so the net field is

E1 + E2 = 2E0 cos ψ sin(ωt + ψ) .


hat the ωt
a.
nents E1 and E2, given by Eqs. 35-20 Phasors that represen
nduring (a )
asors as is discussed in Module 16-6. waves can
and the CHAPTER 1. LIGHT 7 be added
E1 and E2 are represented by phasors find the net wave.
een; theangular speed v. The values
igin at
of the corresponding phasors on the ω
ss and
it over
their projections at an arbitrary E2
regard-
21, the phasor for E1 has a rotation E0
β
q.on35-22;
angle vt !
E 2 f (it is phase-shifted
E
ection
quation E 0 variesωwith
on the vertical axis φ
nctions of Eqs.
E 1 35-20 and 35-21 vary E1
β E0
φ E0
d E2 at any point Pωin
t Fig. 35-10, we ωt
. 35-13b. The magnitude of the vector
es.at35-20
point P, and that wave has aPhasors
cer- that represent (b )
ule 16-6. (a )
E in Fig. 35-13b, we first note that waves
the can
Figure 1.2: be 35-13
Figure added (a)to Phasors representing, at tim
phasors
y are opposite equal-length sides findofthe net wave.
t, the electric field components given by
e) values
that an exterior angle
The intensity of (here f, asfield isEqs.
the electric 35-20 andto35-21.
proportional Both phasors have
its amplitude
sheon theopposite interior angles1(here ω magnitude
two squared,
1
E 0 and rotate with angular
we have
rbitrary E2 I0 = E 2 and speed I = v. Their
E 2 , phase difference is f.
2cµ0 0 2cµ0
E0 (b) Vector addition of the two phasors
rotation so
os b) I E 2
β gives the phasor 1 representing the
-shifted =
I0 E E02
= 4 cos2 ψ , or I = 4I0 cos2 φ .
resultant wave,2 with amplitude E and
φ
sies
1 with
2 f. (35-28)
Using Eq. (1.1.4), we finally phase
obtain the constant
intensity b.
as a function of θ as
21 vary E1
β E 
d

0I(θ) = 4I0 cos2 π sin θ . (1.1.5)
λ
5-10, we ωt
Young’s double-slit experiment does not have practical applica-
e vector tions, but it is the first of many demonstrations that are used to
as a cer- (b ) of some physical entity. For our purposes in
prove the wave nature
that the the rest of the course this is sufficient. Yet interference finds impor-
35-13 (a)although
Phasorsinrepresenting,
Figure
tant applications, different [Link]
these are nice and
sides of t, the electric field components
can be understood fairly simply, we discussgiven bybriefly before we go
them
re f, as onEqs. 35-20
to the and 35-21.
understanding Bothproperties
of other phasors ofhave
light.
es (here magnitude E 0 and rotate with angular
1.1.3
speed Thin-film
v. Their phase interference
difference is f.
(b) Vector
Remember the addition
oil-spill on of
thethe two
road phasors
after rain. On the puddle the oily
place
gives the phasor representing the waves of all wavelengths.
shows nice colors. The sun emits light
These waves are reflected both from the front and back surface of the
resultant wave, with amplitude E and
(35-28) phase constant b.
l
2L ! m , for m ! 0, 1, 2, . . .
n2
(minima — dark film in air),
8 FROM TH I N FI LM S
NTE R FE R E NCE 1065 CHAPTER 1. LIGHT

Before
The interference
Interface depends Interfer
ion can, on the reflections and the
The color
nterface. After path lengths. n1 n2 n3
ge, using light wave
elatively The thick
(a)
waveleng
r2
g. 35-16a ence of th
c
nsmitted Before
r1 Figur
situation After θ b of refract
θ a
index of source. Fo
he wave (b)
i n1 ! n3 in
L
hat is, its perpendic
Figure 35-16 Phase changes when a pulse is or dark t
Figure 35-15 Light waves, represented with
g. 35-16b reflected at the interface between two
ray i, are
Figure incident on a thin film of thick-
1.3: brightly il
ransmit- stretched strings of different linear densi-
ness
ties. The wave speed is greater in the lighter
L and index of refraction n2. Rays r1 The i
entation and r represent light waves that have of the film
string. (a) The incident pulse is in the 2
nusoidal
denser string.
oil film (b) cotes
that The incident pulse been
the water is in reflected
very thinly. Its by thickness
the front and back sur-
is comparable reflected
velength. faces of the film, respectively. (All three
thetolighter string. Only here
the wavelength is therelight,
of visible a phase
i.e. several hundreds of nanometers. the film t
medium change, and only
The light in the
waves reflected
travel our eyes, but there is a path differencetobe-
[Link] are actually nearly perpendicular refraction
e that is the film.) The interference of the waves of
ength.
tween the waves that are reflected from the front and back surfaces. it undergo
r1 and r2 with each other depends on their
action of
For some wavelength the resulting interference is constructive, while
phase difference. The index of refraction
by ray r2, i
for others it is destructive, giving rise to a colorful picture. This phe-
n1 of the medium at the left can differ If the
nomenon can be used to coverfrom objects to reduce or enhance reflection
the index of refraction n3 of the an interfe
of a certain wavelength, thusmediumincrease, or right,
at the decrease thenow
but for transmitted
we are exactl
ones. For instance, thin film onassume that both media are air,reflectivity
a window can enhance the with n1 ! dark to th
in the infrared, thus reduce nthe heating effect of sun light, without
3 ! 1.0, which is less than n2. phase diff
affecting the transmittance for the visible light. The K
Every time light
reflectsoff of a
The physics of thin-film interference is based upon the change of tween the
slow substance,, phase of a reflected light wave. As shown in Fig. 1.3, there are two the path i
there is a pi shift
cases. The wave is reflected from a surface dividing two materials of to b, and t
different indices of refraction. When the wave is coming from the side
through t
1, with smaller index of refraction than on side 2, n1 < n2 , then the
ence betw
phase changes as in the case of reflection of a wave from a fixed end,
between
resulting in a phase shift of 180◦ . If the wave is coming from the side
fference equivalen
1, with greater index of refraction than on side 2, n1 > n2 , then the
phase changes as in the case of reflection of a wave from a free end,
for two re
resulting in zero phase shift. and (2) re
Let us now consider a thin film of width L and refraction index
n n > 1 in air (of refraction index ' 1) as shown in Fig. 1.3, right. The
There is a phase difference between the light reflected from the surface
r2 shown
. Let’s next
5-15. At
CHAPTER 1. LIGHT 9

of 1st incidence (light approaching from the air) and from the 2nd
surface (where the light is approaching from the film). The origin
of this phase difference is two-fold. On the one hand there is 180◦
phase shift between these two waves due to the different relation of
refraction indices: (i) on the first surface n1 = 1 < n2 = n, (ii) while
on the second surface n2 = n > n3 = 1. On the other hand, the wave
travels in the film a path of length of approximately 2L. (Precisely,
it is 2l/ cos β, where β is the angle between the direction of light in
the film and the normal to the surface of the film, but we assume
β is small.) The condition for constructive interference is that the
optical path difference between the two waves results in a phase shift
of 180◦ to compensate for the phase shift due to the different reflection
mechanisms. As the light waves travel with a speed of c/n in the film,
so its wavelength is λ/n in the film. So there will be constructive
interference (maxima) if
 
1 λ
2L = m + , (1.1.6)
2 n
with m being a positive integer. Likewise, the condition for destructive
interference (minima) is
λ
2L = m . (1.1.7)
n
These equations hold if the index of refraction of the film is greater,
or less than the indices of media on both sides of the film because only
in this case there will be a phase shift of 180◦ for reflections at the two
surfaces. In other circumstances, we have to modify the conditions by
taking into account the phase shifts at the different reflections at the
various surfaces.
If the film thickness is not uniform, the conditions for construc-
tive and destructive interference depend on the position and bands
of maxima and minima appear, called fringes of constant thickness.
If such a film is illuminated by white light, one can observe that the
fringes of constant thickness depend on the wavelength and the film
appear in colours.

1.1.4 Single-slit diffraction


The diffraction of waves in a water tank experiment is other char-
acteristics of wave propagation. The same phenomenon can also be
observed if light is let through a single narrow slit. When the width of
10 CHAPTER 1. LIGHT

the slit is comparable to the wavelength of the light, then light beams
flare not only far beyond the geometrical shadow of the slit, but also
show a series of alternating bright and dark bands, similar although
quantitatively different as in the case of the double-slit interference.
Let us consider light waves falling on a barrier with a single slit
on it. In order to describe quantitatively the diffraction pattern we
assume that the distance between the barrier and the screen D is
much larger than the width of the slit d, D >> d. In principle, it
is possible to obtain the diffraction pattern using Fresnel’s principle
without such an assumption, but the required mathematics is more
involved.
In the case D >> d, we can approximate all waves by plane waves,
which can be produced directly using laser beams. The diffraction
pattern of plane waves is called Fraunhofer diffraction. Let us consider
first the point P0 closest to the slit on the screen. No matter from
where in the slit a spherical wave leaves, it arrives on this point with
the same phase, so there will be a constructive interference and a
bright line in the middle point P0 . Let us now consider another point
P1 at distance r from P0 , where we find the first dark line. The
distance to this point from the middle of the slit defines the angle
θ = arctan r/D as shown in Fig. 1.4. If the path from the closest
edge of the slit to P1 differs from the path between the middle of
the slit to P1 by exactly a half wavelength, then the spherical waves
from the edge and the middle arrive at P1 with opposite phase and
interfere destructively. Consider now another pair of spherical waves,
one emerging from a distance x < d/2 from the edge considered before,
and one emerging from a point at the same distance x from the middle,
so the distance between these two points is again d/2. These two waves
arrive at P1 with the same phase difference as the previous two, so
interfere destructively. As x is arbitrary, we conclude that the waves
interfere destructively at P1 giving a dark line. The path difference
between any of these pairs of waves is d/2 sin θ, which has to equal
λ/2 as discussed above, so the condition for the first dark line is

d sin θ1 = λ . (1.1.8)

Note that in our reasoning, the centers of the pairs always lay in
different zones. These two zones called Fraunhofer zones.
We can use similar reasoning to find the condition for the sec-
ond dark line. The only difference compared to the first minimum is
that this time we have to consider four equal Fraunhofer zones (see

[Link]
ximately tocross
reach being
P1with
section longer
a circular edge. than the path traveled by the wavelet of r1. T
this path length difference, we find a point b on ray r2 such that the pa
plifying)
and then from b to P1 matches the path length of ray r1. Then the path length diffe
ncel each tween the [Link]
CHAPTER rays
each
pair of rays cancel
LIGHTis the distance from the center of the slit
other at P1. So 11to b.
point P1. When viewing screen
do all such pairings.C is near screen B, as in Fig. 36-4, the d
n we ex-
pattern on C is difficult to describe mathematically. However, we can sim
y r2 from D
o rays to mathematics considerably if we arrange for the screen separation D to
ays from larger than the Totally
slit width a. Then, as in Fig. 36-5, we can approximate ray
destructive
er of the interference

r2 are in r1 P1
r1
t passing
irst dark r2
a/2 b θ
ifference This path
θ
elet of r2 P0
o display Central axis
b r2 difference
a/2 θ one wave
th length a/2
ence be- Figure 36-5 For D ! a, we Viewing
can θ other, wh
approximate rays r1 and r2screen
as
determine
ffraction being parallel, at
B angle u to the
C
Path length
Incident
mplify the central axis. the interfe
wave difference
be much
Figure 36-4 Waves from the top points of two
r1 and r2 zones of width a/2 undergo fully destructive
FigureC.1.4:
interference at point P1 on viewing screen

Fig. 1.5), and the pairs of waves that cancel at P2 originate from cen-
ength ters separated by distance d/4 on either side of the middle of the slit.
shifts As a result the condition for the second minimum is
rom the
h d sin θ2 = 2λ . (1.1.9)

ence. Using the Fraunhofer zones, it is easy to convince ourselves that the
condition for the mth minimum is

d sin θm = mλ , (1.1.10)

where m is an integer. There are two minima for each m, which


can be accounted for by considering m any integer (both positive and
negative).

1.1.5 Intensity in single-slit diffraction


In between the minima there are maxima approximately half-way,
but not exactly. We now aim at finding the intensity of the light as
a function of the position on the screen characterized by the angle θ
in the case of single-slit diffraction. We use the method of phasors as
r4 difference (a/2)
a/4
θ
P0
(our condition fo
a/4
These rays
a/4 cancel at P2. which gives us
12 CHAPTER 1. LIGHT
1084 CHAPTE R 36 DI FFRACTION B C
Incident
wave Given slit width
(a) fringe above and
D
P2
as being parallel, at angle u to the central axis. We
Narrowing t
gle formed by point b, the top point of the slit, a
while holding th
being
r1 a right triangle,
To see and one of dark
the cancellation, the angles ins
fringes app
path lengthgroupdifference
the rays between
into pairs. rays
andr1theand r2 (w
width of
r1
P1
center of the slit to point b) is thenthe equal to (a/2)
slit width to
r2
θ First
r2 Minimum.
Path length We can repeat this analysis
is 90°. Since the
a/4
nating atdifference
corresponding
between points inthat thebright
two fringe
zone
r3 r1 and r2
a/4 zones) and extending to point P1. EachSecond Min
such pair
r4 difference r(a/2)
3 sin u. Setting thiscentral
common axis as w
path
a/4 into four zones o
θ θ(our condition for the first dark fringe), we have
P0 a/4 r , r , and r from
2 3 4
a/4 r4 ond
a dark fringe l
These rays sin u between
ference ! ,
cancel at P2. 2 2
a/4 which gives us all be equal to l/
θ Path length For D # a, w
B C
a/4 difference between a sin u !
thel central
(firstaxis.
min
Incident r3 and r4
cular line throug
wave Given slit width a and wavelengthries
θ l, Eq. 36-1 tria
of right tel
(a) fringe above and (by symmetry) below We see the centra
from the
(b)
Narrowing the Slit. Note that(a/4) if wesinbegin wi
u. Simila
while
Figure 36-6 (a) holding
Waves from the
thewavelength
top points constant,
r3 and r4we incre
is also (
r1 To see the cancellation,of fourdark
zonesfringes
of width appear;
a/4 undergo that originate
fullyis, the extent of the at diff
corr
group the raysFigure destructive
into pairs.
1.5: interference
and the width of at point P2. (b) For is greater
the pattern) each such
for acase
narrth
D # a, we can approximate rays r1, r2, r3,
and r4the slit parallel,
as being width to the uwavelength
at angle to the (that is, a ! l), t
θ r2
Path length
isaxis.
central 90°. Since the first dark fringes mark the two e
a/4 which givesview
us
difference between that bright fringe must then cover the entire
r1 and r2 Second Minimum. We find the second dar
in the case of double-slit interference, combined with the Fraunhofer
zones. For each rzone
3 we construct the central
phasors, axisand
as we found
the the first darkAll
projection fringes,
Minima ex
of their sum onto the vertical axis into
gives thefour zones
total of equal
electric widths
field as a/4,
a as shown
pattern by in F
splitt
θ
a/4
function r2, r3, andwe
of time. In order to find the intensity, from the
r4 need thetop points
length ofofchoose
the zones to
an even
r4 ond dark fringe above the central [Link]
we hav
the sum of the phasors. low the central a
ference between r1 and r2, that between r2 and r3
When there are many zones, then the all be equal
phasors to l/2.
form an arc of radius
For we cantheapproximate a sin u !
R.a/Each
4
θ phasor on the arc
Path length represent a wavelet D
that# a,
reaches point these four ra
difference between
P on the screen characterized the central axis. To
by the small angle θ. The amplitude display their path You can dif
length rem
r3 and r4
cular line through each adjacent pair one of
in rays, as s
Fig. 36-5,
E(θ) at θthis point is the vector sum of these phasors. At θ = 0, the
ries of right triangles, each of which has a path
ence between th
position of the maximum, the arc becomes Weaseestraight
from thelinetop
and the sum
triangle that the path lengt
of the phasors(b)is the maximal value of (a/4) the electric field, from
sin u. Similarly, Emaxthe. For
bottom triangle, th
θ 6= 0 the length of the arc is
Figure 36-6 (a) Waves from the top points E max , so the central angle of the
r3 and r4 is also (a/4) sin u. Inarcfact,
is the path length
φ of
=fourEmax /R,of so
zones R a/4
width = undergo
Emax /φ. fullyThe phase originate at corresponding points in two adjace
difference between the last
anddestructive interference
first phasor at point
is also φ [Link]
(b) For
in [Link]
1.6. such
Alsocase
fromthethe
pathfigure
lengthwedifference is equal
D # a, we
deduce can approximate rays r1, r2, r3,
that
and r4 as being parallel, at angle u to the a l
sin u ! ,
central axis. 4 2
φ E(θ)which gives us
sin = . (1.1.11)
2 2R a sin u ! 2l (second m

All Minima. We could now continue to loca


pattern by splitting up the slit into more zones o
o diffraction for three values of the ratio a/l.
The wider the slit is, the narrower is the
(36-10) central diffraction maximum.

CHAPTER 1. LIGHT 13
an electromagnetic
ctric field. Here, this
pattern) is propor-
o E u2. Thus,
α α

(36-11) φ
R
1
2 f, we are led to Eq.

e phase difference f Eθ
may be related to a
Em φ Em

Figure 36-9 A construction used to calculate


the intensity in single-slit diffraction. The
es. However, f ! 2a, Figure 1.6: to that of
situation shown corresponds
Fig. 36-7b.
Then we find
E(θ) sin φ
= φ2 , (1.1.12)
Emax 2

and for the intensity


!2
I(θ) E(θ)2 sin φ2
= 2 = φ
. (1.1.13)
Imax Emax 2

Then we use Eq. (1.1.3) where for point P at angle θ, the path differ-
ence between wavelets from adjacent zones of size ∆y is ∆y sin θ, so
the phase difference ∆φ between wavelets from adjacent zones is

∆y
∆φ ≈ 2π sin θ .
λ
The phase difference between the last and the first phasor is the sum of
the phase differences between adjacent phasors over the whole width
of the slit:
d
φ ≈ 2π sin θ ,
λ
14 CHAPTER 1. LIGHT

which leads us to the usual form of I(θ) expressed with the maximum
intensity Imax as
 2
sin α(θ) πd
I(θ) = Imax , α(θ) = sin θ . (1.1.14)
α(θ) λ

1.1.6 Diffraction (or precisely interference) grat-


ings
As the diffraction pattern depends on the wavelength of the light,
and measuring the angles θ is simple, the analysis of diffraction pat-
terns offers a method to measure the wavelength. The only obstacle
to perform a precise measurement is that the bright lines are blurred,
therefore the measurement of their position is not precise. To improve
the position measurement we can use the observation that employing
multiple slits, (i) the bright fringes become narrower, and (ii) faint
secondary lines appear between the primary bright fringes. Further-
more, increasing the number of slits, more secondary lines appear,
but all become fainter. For several thousand slits in one centimeter
these secondary maxima become negligible, while the width δθm of
the primary maxima is inversely proportional to the number of slits
N times the spacing between the slits d,
λ
δθm = , (1.1.15)
N d cos θm
allowing for the precise measurement of the position of the primary
maxima. This observation is used to measure wavelength of light
precisely using diffraction gratings, which has important application
in spectroscopy.
Diffraction gratings usually contain N = 10 000 slits distributed
evenly over a few centimeters, which means the slits are separated by
a few micrometers. We can determine the direction of the mth maxi-
mum θm by the simple requirement that the path difference between
waves originating from neighbouring slits, separated by distance d, is
equal to m times the wavelength (see Fig. 1.7),

d sin θm = mλ , m = 0, ±1, ±2 . . . (1.1.16)

The location of the maxima is independent of N .


an incident wave into its component wavelengths by separat- where it disappears into the
ing and displaying their diffraction maxima. Diffraction by N
(multiple) slits results in maxima (lines) at angles u such that "u hw !
Nd
d sin u ! ml, for m ! 0, 1, 2, . . . (maxima).
CHAPTER 1. LIGHT 15
36-5 DI FFRACTION G RATI NGS 1099

neon laser is shown in Fig. 36-19b. The maxima are now very narrow
ed lines); they are separated by relatively wide dark regions.
P Diffraction Gratings
To point P
on viewing
We use a familiar procedure to find the locations of the bright lines θ One
of the most useful tools in the study of li
screen
screen. We first assume that the screen is far enough from the grat- absorb light is the diffraction grating. This devic
rays reaching a particular point P on the screen are approximately
they leave the grating (Fig. 36-20). Then we apply to each pair of θ arrangement of Fig. 35-10 but has a much great
s the same reasoning we used for double-slit interference. The sep- rulings, perhaps This pathaslength
manydifference
as several thousand pe
een rulings is called the grating spacing. (If N rulings occupy a total between adjacent rays
θ consisting of only five slits is represented in F
d ! w/N.) The path length difference between adjacent rays is determines the interference.
Fig. 36-20), where u is the angle dfrom the central axis of the grating
light isPath sent through the slits, it forms narrow
length
θ analyzed to determine
difference the wavelength of the lig
fraction pattern) to point P. A line will be located at P if the path d
θ between adjacent rays
ce between adjacent rays is an integer number of wavelengths : be opaque surfaces with narrow parallel gro
C Fig. 36-18. Light then scatters back from the gro
sin u ! ml, for m ! 0, 1, 2,λ. . . (maxima — lines), (36-25)
Figure 36-18 An idealized diffraction grating, Figure rather
36-20 Thethan being
rays from transmitted
the rulings in a through open slits
wavelength of the light. Each integer m represents a different line; Pattern.
diffraction grating With
to a distant monochromatic
point
consisting of only five rulings, that produces approximately parallel. The path length dif-
P are light inciden
tegers can be used to label the lines, as in Fig. 36-19. The integers
an interference pattern on a distant viewing gradually increase the
ference between each two adjacent rays is number of slits from two t
the order numbers, and the lines are called the zeroth-order Figure
line 1.7:
ne, with m ! 0), the screen C. line (m ! 1), the second-order line
first-order plot
d sin u, changes
where from
u is measured the typical
as shown. (The double-slit plot of F
rulings extend into and out of the page.)
o on. cated one and then eventually to a simple graph
ng Wavelength. If we rewrite Eq. 36-25 as u ! sin"1(ml/d), we pattern you would see on a viewing screen usin
given diffraction grating, the angle from the central axis to any
third-order line) depends
Exerciseson the wavelength of the light being
Intensity
en light of an unknown wavelength is sent through a diffraction
urements of the angles to the higher-order lines can be Interference
used in ∆θ hw
etermine the wavelength. Even light of several unknown wave-
distinguished and identified in this way. We cannot do that with 3
1. A viewing screen is separated from a double-slit source by 1.2 m.
t arrangement of Module 35-2, even though the same equation
h dependence apply there. The distance
In double-slit between
interference, the two slits is 0.030 mm. The second-
the bright
different wavelengths overlap too much
order to befringe
bright distinguished.
(m = 2) is 4.5 cm from the center line.
θ De-0°
termine the wavelength of the Figure
light!Figure 36-19 (a) The intensity plot produced
nes 36-21 The half-width #uhw of the cen-
by a diffraction grating with a great many
tral line is measured from the center of that
lity to resolve (separate)[Link]
Aofpossible
different wavelengths
means for depends on an
making airplane invisible to radar ishere
to labeled
he lines. We shall here derive an expression for the half-width of line torulings consists
the adjacent of narrow
minimum on a plotpeaks,
of
coat the plane with
e (the line for which m ! 0) and then state an expression for the
an antireflective u polymer.
like Fig. 36-19a. If radar
with their order numbers m. (b) The
I versus waves
havethea half-width
the higher-order lines. We define wavelength = 3.00 cmcorresponding
of λ line
of the central and the index brightof fringes seen on the
refraction
ngle #uhw from the center of ofthethe linepolymer
at u ! 0 outward
is n = to 1.50,
where how thick screen (d) are called
would lines
youandmakeare here also
the 3
vely ends and darkness effectively begins with the first minimum labeled with order numbers m.
coating?
such a minimum, the N rays from the N slits of the grating cancel
The actual width of the central line is, of course, 2(#uhw), but line Top ray
To first
3. Solar cell devices that generate electricity when exposed
ally compared via half-widths.) minimum
to sun-
36-1 we were also concerned light
with theare often coated
cancellation of a great with
many a transparent, thin film of silicon
to diffraction through a single slit. We obtained
monoxide (SiO, Eq.n= 36-3, which,
1.45) to minimize ∆θ hwreflective Bottomlosses from the
e similarity of the two situations, we can use to find the first ray
surface. Suppose that a silicon solar Nd cell (n = 3.5) is coated with
. It tells us that the first minimum occurs where the path length
ween the top and bottom rays a thin
equalsfilm
l. Forof silicondiffraction,
single-slit monoxide for this purpose. Determine the
Path length
is a sin u. For a grating of N rulings, each separated from the next ∆θ hw difference
minimum film thickness
he distance between the top and bottom rulings is Nd (Fig. 36-22),
d that produces the least reflection at
a wavelength of 550
th length difference between the top and bottom rays here is nm, near the center of the visible spectrum.
hus, the first minimum occurs where Figure 36-22 The top and bottom rulings of
4. A thin film of oil (n = 1.25) isa diffractionlocatedgrating
on ofa Nsmooth
rulings are wet pave-
Nd sin #uhw ! l. (36-26) separated to
by Nd. The pavement,
top and bottom rays
ment. When viewed perpendicular the the film
is small, sin #uhw ! #uhw (in radian measure). Substituting this in passing through these rulings have a path
reflects
the half-width of the central line as
most strongly red light at λ = 640 nm and
length difference of Nd sin #uhw, where
r reflects no
blue light at λ = 512 nm. How is the angle
thick
#uhw is to theoil
the firstfilm?
minimum.
l b
#u hw ! (half-width of central line). (36-27) (The angle is here greatly exaggerated
Nd for clarity.)
16 CHAPTER 1. LIGHT

Diffraction

1. A slit of width d is illuminated by white light. The first mini-


mum for red light (λ = 628 nm) falls at θ = 15◦ . What is the
wavelength of the light whose first diffraction maximum (except
the central one) falls also at 15◦ , thus coinciding with the first
minimum for red light?
2. Calculate, the relative intensities of the successive secondary
maxima in a Fraunhofer diffraction pattern.
3. We measure the wavelength of an unknown monochromatic light-
source with a diffraction grating with 3000 rulings per cm. We
find the first maximum at α1 = 10.37◦ . (a) What is the wave-
length of this light?
(b) Where can we find the second maximum?

4. Monochromatic light from a helium-neon laser (λ = 632.8 nm)


is incident normally on a diffraction grating containing 6 000
grooves per centimeter. Find the angles at which the first- and
second-order maxima are observed!
5. A helium-neon laser (λ = 632.8 nm) is used to calibrate a diffrac-
tion grating. If the first-order maximum occurs at 20.5◦ , what
is the spacing between adjacent grooves in the grating?
6. We analyse the wavelength of the light emitted by a source using
a diffraction grating with 2000 rulings/cm. The angle separation
between the central and first maxima is θ = 6.765◦ . What is
the possible source of this light?
CHAPTER 1. LIGHT 17

Polarization

1. Draw the amplitude of the wave A cos(ωt − kz)i + A cos(ωt −


kz +π/3)j in the polarisation plane for {A} = 1 and time period
T = 20 s. What kind of polarisation is this?
2. What is Brewster’s angle for glass (with refraction index 1.5) in
air? What is Brewster’s angle for the same glass in water (with
refraction index 1.33)?
3. If the interplanar spacing of NaCl is d = 281 pm, what is the
predicted angle θ at which λ = 0.140 nm X-rays are diffracted
in a first order maximum?

1.2 Blackbody radiation


We now proceed to understand light better and consider one of the
major problems at the end of the 19th century: the thermal radiation
emitted by an ideal blackbody radiator. We know from experience
that heated objects emit electromagnetic radiation, which can even
be in the form of light if the temperature of the object is high enough.
The emitted radiation of a blackbody radiator depends only on its
temperature and not on its material, the nature of its surface, or
anything else. A good model of a blackbody radiator is a small hole on
a piece of metal kept at uniform temperature and containing a cavity
inside. The atoms on the inner wall of the cavity oscillate as they have
thermal energy. This oscillation makes them emit electromagnetic
waves that fills the cavity. In order to sample that internal radiation,
we measure the radiation that escapes through the hole that should
be small not to alter the radiation inside the cavity. We measure how
the intensity of the radiation depends on wavelength (or equivalently,
on frequency).
The spectral radiancy is defined as
dI
S(λ) = (1.2.1)

where the intensity I is the emitted power over a unit area of the
radiator (in our example the hole) into a unit of solid angle. Thus [S]
= W/m3 , i.e. the physical dimension of the spectral radiancy is the
18 CHAPTER 1. LIGHT
38-4 TH E B I RTH OF QUANTU M PHYSICS 1165

ctric effect and Compton 50


tum physics, let’s back up
Spectral radiancy (W/cm2 " mm)

uantized energies gradu- Experiment T ! 2000 K


he story begins with what 40
h was a fixation point for
hermal radiation emitted 30
a radiator whose emitted Classical
theory
e and not on the material
urface, or anything other 20
e was the trouble: the
m the theoretical predic- 10

n ideal radiator by form-


e cavity walls at a uniform 0 1 2 3 4 5 6
of the body oscillate (they Wavelength (mm)
to emit electromagnetic Figure 38-8 The solid curve shows the experimental spectral ra-
hat internal radiation, we diancy for a cavity at 2000 K. Note the failure of the classical
some of the radiation can theory, which is shown as a dashed curve. The range of visible
alter the radiation inside wavelengths is indicated.
intensity of the radiation
Figure 1.8:
d by defining a spectral
given wavelength l:
power
. (38-12)
same as that of power density, so it must be proportional to the energy
area
"! unit
mitter wavelength "
density
gth range dl, we have the intensity (that
of the radiation inside the cavity, the source of the radiated
power.
n the wall) that is being emitted in the We can imagine the radiation inside the cavity as standing

e experimental results for aelectromagnetic


cavity with a waves. Assuming two degrees of freedom for the two
f wavelengths. Although such a radiator
can tell from the figure thatpolarisation
only a small states of the wave (the electric and the magnetic field),
in the visible range (which is colorfully
the radiated energy lies inthe equipartition theorem suggests an average energy of a single wave
the infrared

hεi =
physics for the spectral radiancy, for a kT where T is the temperature inside the cavity. Then on
dimensional grounds we expect that the spectral radiancy is
sical radiation law), (38-13)

9-7) with the value ckT


! 8.62 " 10#5 eV/K. S(λ) ∝ , (1.2.2)
for T ! 2000 K. Although the theoreti- λ4
t long wavelengths (off the graph to the
which is supported by measurements only for large values of the wave-
rt wavelength region. Indeed, the theo-
e a maximum as seen in the measured
length (see Fig. 1.8). It fails clearly for small λ because it predicts a
inity (which was quite disturbing, even

devised a formula for S(l) diverging


that neatly spectral radiancy, in contradiction not only to the data, but
elengths and for all temperatures:
also because it would lead to infinite radiated energy upon integra-
(Planck’s radiation law). (38-14)
#1 tion over λ. In fact, at the end of the 19th century, the total radiated
power was known to be proportional to the fourth power of tempera-
ture according to a phenomenological law of Josef Stefan,

P = σAT 4 (1.2.3)

–now known as Stefan-Boltzmann law because Ludwig Boltzmann was


the first to derive it from theoretical considerations as follows.
The SI unit of energy density is J/m3 = N/m2 , which is the unit
of pressure, so pressure p and energy density ε in a substance must be
proportional, p = wε, which is called the equation of state. In fact,
we know that p = kT for gas and there are three degrees of freedom
A "relativistic gas" typically refers to a gas composed of particles (such as electrons or photons) that are moving at speeds comparable to the speed of light (
c). In such gases, the effects of special relativity, as described by Albert Einstein's theory of special relativity, become significant. This is in contrast to a "
non-relativistic gas," where the speeds of particles are much lower than
c, and classical physics can be used to describe their behavior.

CHAPTER 1. LIGHT 19

for a molecule in a gas in ideal-gas approximation, hence in this case


w = 2/3. In the case of a relativistic gas there are six degrees of
freedom for each “molecule” because each particle has two polarization
states, hence w = 1/3. From the 1st law of thermodynamics dU =
T dS − pdV , so    
∂U ∂S
=T − p. (1.2.4)
∂V T ∂V T
Differentiating with respect to T while keeping V fixed, we obtain
Polarization States: When considering photons (particles
Maxwell’s relation of light) in a gas, they have an additional set of degrees of
    freedom associated with their polarization states. Photons
∂S ∂p are electromagnetic waves, and their electric and
0= − .
magnetic fields can oscillate in various directions
∂V T ∂T V perpendicular to their direction of propagation. The
number of polarization states for a photon is determined
Substituting U = εV into the left hand side and Maxwell’s relation
into the right hand side of Eq. (1.2.4), we obtain
by the number of independent directions in which its
  electric field can oscillate. In three-dimensional space,
∂p there are two independent polarization states for photons,
ε=T − p . typically denoted as vertical and horizontal polarizations.
∂T V

Now using the equation of state for radiation, p = ε/3, we find

4ε 1 dε
=
3T 3 dT
where we could replace the partial derivative with ordinary derivative
as it depends only on the ratio ε/T . Integration with respect to T
leads to the Stefan-Boltzmann law in Eq. (1.2.3).
In Eq. (1.2.3) A denotes the area of the radiator and

W
σ = 5.670373 · 10−8 (1.2.5)
m2 K4
is the Stefan-Boltzmann constant that can be measured, but we shall
see that it is related to fundamental constants of nature.
The measured spectral radiancy has a maximum at λ = λmax .
In 1893 Wilhelm Wien published a theoretical article where he ar-
gued that the product λmax T (more precisely, the ratio T /fmax ) is a
constant,
λmax T = 2.8977729(17) mm · K , (1.2.6)
which is called Wien’s displacement law, demonstrated on Fig. 1.8,
right.
20 CHAPTER 1. LIGHT

A solution to understand the form of the spectral radiancy function


was given by Max Planck in 1900. He wrote down his formula,

2c2 h 1
S(λ) = 5 hc/λkT
, (1.2.7)
λ e −1
whose origin was not known until in 1917 Einstein produced a deriva-
tion assuming energy quanta both for the atoms of the cavity wall
as in the Einstein model of solids and for the electromagnetic radia-
tion. For the former we already found in the Thermodynamics course
that the Boltzmann distribution with quantized energy levels of the
Discrete Energy Levels: Electrons in an atom, vibrational modes of a
oscillators predicts molecule, energy levels of a quantum harmonic oscillator, quantum
n classical statistical mechanics, the equipartition theorem states that each ε0
degree of freedom (kin+pot) contributes hεi = ε0 ,states of a particle in a potential well.(1.2.8)
kT/2 to the average energy of the system. This theorem works well for exp ( kT ) − 1
systems where energy levels are continuous, and classical physics is
applicable as average energy per oscillator instead of kT . Thus, instead of
Eq. (1.2.2) we should have Continuous Energy Levels: Kinetic energy of a classical gas particle,
kinetic energy of a macroscopic object, energy levels of a classical
c ε0 harmonic oscillator, classical states of a macroscopic system.
S(λ) ∝ 4 ε0 . (1.2.9)
λ exp ( kT ) − 1

For the quantum of the electromagnetic radiation Einstein assumed


ε0 = hf where h is Planck’s constant. As f = c/λ, Eq. (1.2.9) leads to
formula (1.2.7) (with a factor of 2 as constant of proportionality). We
can also express the spectral radiancy as a function of frequency, which
hf
simplifies the exponent of the exponential function to kT . However,
we have to keep in mind the definition of S(λ) as given in Eq. (1.2.1),
i.e.
dI dλ 2c2 hf 5 1 c
S(f ) = = S(λ) =
df df λ→c/f c5 ehf /kT − 1 f 2
(1.2.10)
2hf 3 1
= 2 hf /kT .
c e −1
If the radiator is not a perfect black body, then the spectral radi-
ancy in Eq. (1.2.7) is reduced by a factor that depends on the wave-
length (λ) ∈ [0, 1].
We can use Eq. (1.2.10) to compute the total radiated power. We
have to integrate (a) over the area A of the radiator, (b) over the half
of the complete solid angle, but taking into account that at an angle θ
away from the direction perpendicular to the surface, the effective area
is on A cos θ as shown in Fig. 1.9 and (c) over all possible frequencies.
The first integral in (a) is easy as the integrand does not depend on
CHAPTER 1. LIGHT 21

the surface element, so it simply gives an overall factor of A. The


second integral is also easy if we notice that the integrand depends
on the angle only as shown in Fig. 1.9, so Effective
it gives anarea seen
overall at of
factor an angle θ

Z 2π Z π/2 Z 1
dφ cos θ sin θdθ = 2π xdx = π .
0 0 0

To perform the third integral in


(c), it is convenient to change the
variable of integration f to the
hf
dimensionless variable x = kT ,
with Jacobian kT /h, so we ob-
tain

(kT )4
P = 2πcA I (1.2.11) dA cosθ
(hc)3
dA
where the integral

x3 dx
Z
I= Figure 1.9:
0 ex − 1

is a number. The computation


of this number is not easy, but
can be looked up. We find that I = 6ζ(4) where ζ(4) = π 4 /90 is the
value of Riemann’s ζ(x) function at x = 4. Comparing Eqs. (1.2.3)
and (1.2.11), we find the value of the Stefan-Boltzmann constant ex-
pressed in terms of fundamental constants,

2π 5 k 4
σ=
15 c2 h3
which gives the numerical value quoted in Eq. (1.2.5).
From the analytic form of the spectral radiancy function S(λ), we
can easily find the position of its maximum λmax from the condition

dS
= 0.
dλ λ=λmax

Again, it is more convenient to use the dimensionless variable x =


hc/λkT . As λ = hc/xkT and dS/dλ = (dS/dx)(dx/dλ), we have to
22 CHAPTER 1. LIGHT

find the solution to the equation

d x5 5x6 x7 e x
x2 = − =
dx ex − 1 ex − 1 (ex − 1)2
5ex − xex − 5
= x6 = 0,
(ex − 1)2

which we can do numerically. We find x = 4.96511, and we can express


the numerical value in Eq. (1.2.6) with fundamental constants as

hc
λmax T = .
4.96511k
As discussed above the spectral radiancy results from the electro-
magnetic radiation inside the cavity. The function S(f ) has SI unit
W s/m2 ,1 so the differential energy density (with unit J/m3 ) in the
frequency range [f, f + df ] of that radiation should have the same
form,
S(f )
dε(f, T ) = 4π df (1.2.12)
c
where the factor of 4π is due to the integration of the spectral radiancy
over the complete solid angle (imagine that the radiator is the Sun).
In order to find the total energy density inside the cavity (or Sun), it
hf
is again more convenient to use the dimensionless variable x = kT as
integration variable instead of f , so

(kT )4 x3 dx
dε(x, T ) = 8π . (1.2.13)
(hc)3 ex − 1

Exercises
Black-body radiation

1. The surface temperature of our Sun is about T = 5500 K, while


its diameter and mass are d = 1.4 · 109 m and m = 2.0 · 1030 kg.
We may assume that it radiates electromagnetic radiation as a
black-body radiator. Compute the absolute and relative losses
of mass in one second due to this radiation.
2. Compute the total radiated power from a black radiator at tem-
perature 45.0◦ of surface area 2.00 m2 .
1 Notice that the definitions of S(λ) and S(f ) are only similar, not identical,

hence the difference in their units.


CHAPTER 1. LIGHT 23

3. Compute the wavelength of electromagnetic radiation where the


radiated power from a black radiator at temperature 45.1◦ is
maximal.
4. What is the value of the radiated power in the wavelength range
of ∆λ = 100 nm from a black radiator of surface area 2.00 m2 at
temperature 45.0◦ at the peak of the spectrum?
5. Compute the energy density of the Cosmic Microwave Back-
ground radiation (CMBR) of temperature T = 2, 726 K.
6. Compute the density of the photons in the CMBR.
7. What fraction of the photons have energy greater than E =
2.2 MeV in a photon gas of temperature T = 900 MK?

1.3 Particle properties of light


The great success of the quantised energy hypothesis in explaining the
spectrum of black-body radiation gives a strong support to imagining
light as flow of particles. There is another experiment, whose inter-
pretation was also given by Einstein in 1905, which gave boost to the
same idea.

1.3.1 Photoelectric effect


The experimental setup is shown in 1156 CHAPTE R 38 PHOTONS AN D MATTE R W
Fig. 1.10. It consists of a vacuum tube
with a target T inside this tube (made Quartz
window
The Photoele
Vacuum
of clean metal and called cathode) and If you direct a
C
a collector cup C (also metal and called surface, the ligh
electrons from
anode). Through the quartz window T Incident
light including camco
at the end of the tube we direct light i Let us anal
A of Fig. 38-1, in
onto the target, with variable frequency i
electrons from
f and intensity I. A variable voltage V collector cup C
V is maintained between the cathode Sliding contact lection produce

and the anode. The current between – + – + First Photoelectr


the cathode and the anode (if any) is We adjust the p
measured with meter A. that collector C
Figure 38-1 An apparatus used to study the ference acts to
There are two steps of the photo- photoelectric effect. The incident light a certain value,
shines on target T, ejecting electrons, which
electric experiment. In the first one we are collected by collector cup C. The elec-
meter A has ju
Figure 1.10: electrons are t
trons move in the circuit in a direction op-
posite the conventional current arrows. The kinetic energy o
batteries and the variable resistor are used
to produce and adjust the electric potential
24 CHAPTER 1. LIGHT

adjust the frequency of the light such


that we can measure current at V = 0.
This current is due to the electrons ejected by the illuminated cath-
ode and collected by the anode. Then we vary the voltage V until
it reaches a value Vstop when the current vanishes. At V = Vstop
the most energetic electrons, ejected by the incident light from the
cathode, are stopped by the electric field and turned back just before
reaching the anode. This way we can measure the kinetic energy of
the most energetic electrons using the equation

(max)
Ek = eVstop (1.3.1)

where e is the unit charge. Varying the intensity of the incident light,
(max)
we find that the stopping voltage, hence Ek does not depend on I.
In a classical picture, the alternating electric field in the light makes
the electrons oscillate inside the target. If the amplitude, hence the
intensity of the electric field is large enough it makes the electron
oscillate with so large amplitude that it can break free. So classically
(max)
we would expect that with increasing intensity Ek increases, but
this is not what happens. The intensity of the light does not influence
the maximum kick it gives to the ejected electrons.
If we think of the incident light as flow of particles with definite
energy that depends only on the frequency, then the energy that can
be transfered to the electron from the light is that of a single light
particle, called photon. Increasing the intensity increases the number
of photons in the light beam, but not the energy of the photons as
that depends only on the frequency, E = hf . Thus, the frequency
38-2 TH E PHOTOE LECTR IC E FF
being fixed, the energy transferred to the electrons is also fixed.
In the second experiment we mea- Electrons can escape only The escaping electron’s
if the light frequency kinetic energy is greater
sure the stopping voltage as a function exceeds a certain value. for a greater light frequency.

of the frequency. As shown in Fig. 1.11, Visible Ultraviolet

we find that they a linearly proportional


Stopping potential Vstop (V)

above a certain cutoff frequency fc , 3.0

a
2.0
Figure 38-2 The stopping po-
Vstop = a(f − fc )Θ(f − f ) , (1.3.2)
c Vstop as a function of
tential
the frequency f of the inci- 1.0 Cutoff
dent light for a sodium target frequency f 0 c b
T in the apparatus of Fig. 38-
independently of the intensity. (Θ(x) = 1. (Data reported by R. A.
0
2 4 6 8 10 12
Millikan in 1916.) Frequency of incident light f (10 Hz) 14
1 if x > 0 and 0 if x < 0, called the step
function.) Based on classical physics
The existence ofwe a cutoff frequency is, however, just what we should expect
if the energy is transferred via photons. The electrons within the target are held
expect that if the intensity is sufficiently
there by electric forces. (If they weren’t, they would drip out of the target due to
Figure
the gravitational force on them.) To just escape from the 1.11:
target, an electron must
pick up a certain minimum energy !, where ! is a property of the target material
called its work function. If the energy hf transferred to an electron by a photon
exceeds the work function of the material (if hf " !), the electron can escape
the target. If the energy transferred does not exceed the work function (that is,
if hf # !), the electron cannot escape. This is what Fig. 38-2 shows.
CHAPTER 1. LIGHT 25

high, then electrons can be ejected in-


dependently of the frequency, contrary
to the observation.
Thinking in terms of photons, we expect exactly the behaviour
seen in Fig. 1.11. The electrons in the metal target are kept inside by
an average attractive electric field produced by the positively charged
atoms. In order that an electron breaks free from this potential valley,
some minimum energy Wmin has to be transferred to it. Wmin is
characteristic to the type of the target, and called work function. If
the energy of the photons is larger than the work function, hf >
Wmin than the electron will be kicked out, if it is smaller, the electron
remains inside.
It was Einstein who summarized these observations into the pho-
toelectric equation
(max)
hf = Wmin + Ek , (1.3.3)
and was awarded the Nobel prize for this achievement. Eq. (1.3.3)
simply expresses the conservation of energy in the elementary scat-
tering process between the photon of the illuminating light and the
electron of the target, namely the energy of the photon is partly
used to kick out the electron–this is Wmin –and the remaining energy
(max)
Ek = hf − Wmin is passed to the electron in the form of kinetic en-
ergy. Combining Eqs. (1.3.1) and (1.3.3) we can express the stopping
voltage as a function of the frequency as
 
h Wmin
Vstop = f− . (1.3.4)
e h
Comparing Eq. (1.3.4) to Eq. (1.3.2) we see that we can measure the
ratio h/e from the plot of stopping voltage versus frequency. For
instance, from Fig. 1.11 we find
h
' 4.1 · 10−15 V · s .
e
Using the value for the unit charge, e ' 1.6 · 10−19 C, we can deduce
the value of Planck’s constant,

h ' 6.6 · 10−34 J · s .

We also find that the work function and the cutting frequency are
related by
Wmin = hfc . (1.3.5)
26 CHAPTER 1. LIGHT
38-3 PHOTONS, M OM E NTU M , COM PTON SCATTE R I NG, LIG HT I NTE R FE R E NCE 1159
1.3.2 Compton scattering
Ideas
hough it is massless,The explanation
a photon has momentum,of the photoelectric effect using h the concept of light
h is related to its energy E, frequency f, and "l ! (1 # cos f),
particles was very encouraging. To extend this mc hypothesis further, in
length by where m is the mass of the target electron and f is the angle
1916 Einstein proposed that photons also have momentum. As the
at which the photon is scattered from its initial travel direction.
hf h
unit !of the
p! . ratio energy/momentum
● Photons: When is light
the interacts
unit of withspeed and
matter, the in theis
interaction
c l
case of light the only quantity particle-like,
that occurring
can have at a point
unitandoftransferring
speed isenergyc, itandis
Compton scattering, x rays scatter as particles (as momentum.
ons) from loosely boundnatural
electronstoin aassume
target. that a photon
● Wave: of Whenenergy
a singleE = hf
photon has momentum
is emitted by a source, we
he scattering, an x-ray photon loses energy and interpret its travel as being that of a probability wave.
entum to the target electron. ● Wave:E When many
h photons are emitted or absorbed by
e resulting increase (Compton shift) in the photon matter, = the combined light as a classical(1.3.6)
p = we interpret electro-
length is magneticcwave. λ

where λ = c/f is the wavelength of the photon.


The proof for the concept of pho-
tons Have Momentum ton momentum was given experimen-
16, Einstein extended his concept
tally by Arthurof light Compton
quanta (photons) by proposing
in 1923. He Detector
a quantum of light has linear momentum. For a photon with energy hf, the
nitude of that momentum directed
is a beam of X-rays of definite
wavelength λ = 71.1 pm on a carbon Incident
x rays
hf h λ'
ptarget
! ! and(photon
measured
momentum),the wavelengths (38-7) λ
Scattered
c % φ x rays
and intensities of the scattered X-rays T
e we have substituted as for f from Eq. of
a function 38-1 (f !
the c/l). Thus, when
scattering angle a photon
φ as Collimating
acts with matter, energy and momentum are transferred, as if there were slits
shown
ision between the photon and in Fig.
matter in 1.12. The
the classical results
sense at four
(as in Chapter 9). Figure 38-3 Compton’s apparatus. A beam
n 1923, Arthur Compton different scattering angles are shown in of x rays of wavelength l ! 71.1 pm is
at Washington University in St. Louis showed that
momentum and energy are transferred via photons. He directed a beam of x
Fig. 1.13. The interesting feature is that directed
of wavelength l onto a target made of carbon, as shown in Fig. 38-3. An x ray
onto a carbon target T. The x rays
scattered from the target are observed at
Figure 1.12:
there are two peaks, separated
orm of electromagnetic radiation, at high frequency and thus small wave- by a dif- various angles f to the direction of the inci-
h. Compton measured the wavelengths
ference in wavelengthand intensities
∆λ(φ) of the=x λ rays − λ dent
0 that beam. The detector measures both the
intensity of the scattered x rays and their
scattered in various directions from his carbon target.
that depends on the scattering
Figure 38-4 shows his results. Although there is only a single wavelength angle. You are challenged to find out
wavelength.
71.1 pm) in the incident
whatx-ray youbeam,wouldwe see that thein
expect scattered x rays con-
a classical theory. Instead, we immediately
a range of wavelengths with two prominent intensity peaks. One peak is
try to employ the photon hypotheses and see if we can understand the
ered about the incident wavelength l, the other about a wavelength l$ that
nger than l by an amountobserved
"l, which pattern
is calledofthethe intensity
Compton shift. Thedistribution
value of the scattered light.
e Compton shift varies with Wetheuse angletheat which the scattered xof
conservation rays are de- and momentum to find the
energy
d and is greater for a greater angle.
relation
Figure 38-4 is still another puzzle between the wavelength
for classical physics. shift and scattering angle. We use
Classically, the incident
beam is a sinusoidallytheoscillating electromagneticnotation
four-component wave. An electron
for ainunified
the treatment of energy and
momentum. Assuming that the electron is at rest before the collision,
φ = 0°
we can choose theφ =z45°axis of our coordinate system aligned φwith
φ = 90° = 135°
the
direction of the incoming photon. Then the four-momenta of the
Intensity

Intensity
Intensity

photon γ and the electron e before collision can be parametrized as


  ∆λ

∆λ
h ∆λ 2
pγ = hf, 0, 0, , pe = (mc , 0, 0, 0) (1.3.7)
70 75 70 75 λ 70 75 70 75
Wavelength (pm) Wavelength (pm) Wavelength (pm) Wavelength (pm)

38-4 Compton’s resultswhere m is


for four values the
of the massangle
scattering of [Link]
Note electron.
that the The momenta of the particle
pton shift "l increases as the scattering angle increases.
ntum), length. (38-7)
Compton measured the wavelengths and intensities of the x ra
Scattered
λ
were scattered in various directions fromφ his xcarbon rays
target.
T
Figure 38-4 shows his results. Although there is only a single wave
! c/l). Thus, when a photon
(l ! 71.1 pm) in the incident
Collimatingx-ray beam, we see that the27 scattered x ra
ransferred, as if there 1.
CHAPTER wereLIGHT
slits with two prominent intensity peaks. One
tain a range of wavelengths
sical sense (as in Chapter 9). 38-3 Compton’s apparatus. A beam
centered about theFigureincident wavelength l, the other about a wavelength
sity in St. Louis showed that of x rays of wavelength l ! 71.1 pmthe
is Compton shift. Th
is longer than l by an amount "l, which is called
ons. He directed a beam of x directed
of the Compton shift variesonto
witha carbon target
the angle atT. The x the
which raysscattered x rays
s shown in [Link]
38-3. An x ray scattered from the target are observed at
and is greater for a greater angle.
uency and thus small wave-
Figure
various angles f to the direction of the inci-
38-4 is still another puzzle for classical physics. Classically, the i
ntensities of the x rays that dent beam. The detector measures both the
x-ray beam is a sinusoidally
intensity ofoscillating electromagnetic
the scattered x rays and their wave. An electron
arget.
is only a single wavelength wavelength.
hat the scattered x rays con-
ntensity peaks. One peak is φ = 0° φ = 45°

r about a wavelength l$ that

Intensity
Intensity

Intensity
he Compton shift. The value
the scattered x rays are de-

ysics. Classically, the incident ∆λ


etic wave. An electron70in the 75 70 75 7
Wavelength (pm) Wavelength (pm) W

Figure 38-4 Compton’s results for four values of the scattering angle f. Note that the
= 45° φ = 90° as the scattering angleφincreases.
Compton shift "l increases = 135°
Intensity

Intensity

∆λ

∆λ

5 70 75 70 75
m) Wavelength (pm) Wavelength (pm)

ng angle f. Note that the


s.
Figure 1.13:
28 CHAPTER 1. LIGHT

after collision can be written as


 
0 h h
qγ = hf , 0 sin φ, 0, 0 cos φ ,
λ λ
  (1.3.8)
0 2 h h h
qe = h(f − f ) + mc , − 0 sin φ, 0, − 0 cos φ
λ λ λ
where we have built the conservation of energy and momentum into
this form of the four-momenta. The components of the four momen-
tum for the electron have to fulfill the dispersion relation, so
 2  2  2
hc hc hc 2 hc hc
− 0 + mc2 − sin φ − − cos φ = m2 c4 .
λ λ λ0 λ λ0
After expanding the squares, we obtain
1 1 1 mc 1 mc 1 1
− 0
+ − 0 + cos φ = 0 ,
λλ λ h λ h λ λ0
or after multiplying with λλ0 h/(mc),
h
λ0 − λ = (1 − cos φ) .
mc
This shift–called Compton shift–agrees with the difference between
the positions of the two peaks in Compton’s experiment. The peaks
have some width because the hit electrons are not at rest, and their
initial motion screens the infinitely narrow peak. The factor h/(mc) '
2.43 pm is called Compton wavelength of the electron. If we were
experimenting with other charge particle X, then the factor h/(mX c)
would give the Compton wavelength of that particle.
The peak at the initial wavelength has a different origin. It is
due to the scattering of the photons off electrons tightly bound to
the nuclei of the carbon atoms. Thus the elastic scattering occurs
effectively between the photon and the whole carbon atom. As the
mass of the whole atom is about 22 thousand times larger than that of
the electron, the corresponding Compton wavelength is about 0.11 fm,
so the Compton shift is so small that it is not observed.

Exercises
Photons
1. Using the plots for the stopping voltage versus frequency for
Cesium, Potassium, Sodium and Lithium in Fig. 1.14, estimate
the work function for these metals.
hown in countless other experiments, light is in fact quantized as
nstein’s explanation of the photoelectric effect is not the best ar-
fact. CHAPTER 1. LIGHT 29

um
m,

m
um
m
Vstop

iu
siu

ssi

di

th
e targets

ta
Ce

So

Li
Po
ots accord-
5.0 5.2 5.4 5.6 5.8 6.0 6.2
14
f (10 Hz)

Figure 1.14:

2. We use light of wavelength λ = 400 nm in a photoelectric exper-


iment and we find that the kinetic energy of the most energetic
nd work function ejected electrons is exactly half of the energy of the irradiating
photons. The size of the target cathode is A = 50 mm2 , the in-
tensityFrom
Calculations: of the that
light last
is I =idea,
1 mW/m Eq.2 38-5 then
and one outgives us, with
of a thousand
f ! f0, photons ejects an electron. (i) Compute the work function in the
cathode. (ii) Compute the maximal kinetic energy of an elec-
tron if the same 0!0%
hfcathode & ! &. with light of wavelength
is illuminated
λ0 = 350 nm (iii) Compute the number of electrons emitted by
In Fig. 38-2,
thethe cutoff
cathode frequency
in one second using f0 islight
theoffrequency at which
both cases separately.
y
the plotted line intercepts the horizontal frequency axis,
(iv) Compute the power in the electron beam that leaves the
: cathode.
about 5.5 " 10 14 Hz. We then have
5 3. We use X-ray of energy E = 51 keV in a Compton scattering
experiment. We find that#34 the wavelength of the 14 X-ray increases
a &! hf ! (6.63 " 10 J$s)(5.5 " 10 Hz)
by 100% in a certain direction of the scattered light. (i) What
h is the initial and scattered wavelength of the photons in this
! 3.6 " 10 #19
experiment? J ! is2.3
(ii) What theeV. (Answer)
direction of the scattered photon?
(iii) What is the increase of the energy of the electron assuming
that it was at rest initially? (iv) In what direction will move
the kicked electron with respect to the direction of the incoming
available at WileyPLUS
X-ray beam?
4. Compute the energy of the photon whose wavelength is equal to
the Compton wavelength of the electron.
5. Consider an X-ray beam of wavelength λX = 100 pm and a γ-ray
beam from 137 Cs atoms of wavelength λγ = 1.88 pm. We view
the radiation scattered off free electrons at an angle θ = 90◦
PTON SCATTERING, LIGHT INTERFERENCE
with respect to the incident beam.
(i) Compute the Compton-shift for both beams.
30 CHAPTER 1. LIGHT

(ii) Compute the kinetic energy given to the recoiling electron


in bot cases.
(iii) What percentage of the incident photon energy is trans-
ferred to the electron in the two cases?
Chapter 2

Particles

2.1 Wave properties of matter


In the previous chapter we have found that light shows the char-
acteristics of waves as well as particles. A beam of light produces
interference effect, but it transfers energy and momentum to matter
only at points, via photons. In 1924, Luis de Broglie suggested that
a beam of electrons, that had been thought of particles of definite
energy and momentum, could also thought of a wave, called matter
wave. He assumed that the momentum-wavelength relation for the
photon, p = h/λ, could also be applied to electrons. However, in this
case we can measure the momentum of the electron more easily, so
he reversed the relation to find the wavelength of the matter wave as
λ = h/p, which is called de Broglie wavelength of the moving particle.
In order to gain some quantitative experience about the wave-
lengths of matter waves, let us compute the de Broglie wavelength of
electrons being accelerated over a potential difference of 100 V. As-
suming that the electron started from rest, its kinetic energy after the
acceleration is 100 eV, which is much smaller than its rest energy of
511 keV. Thus we may approximate the momentum of the electron in
the non-relativistic limit, i.e.
p p
p = 2mEk ' 2(9.11 · 10−31 kg)(100 eV)(1.60 · 10−19 J/eV)
' 1.8 · 10−24 kg · m/s ,
so the corresponding wavelength is λ ' 370 pm, which is about a
thousand times shorter than that of visible (purple) light.
31
32 CHAPTER 2. PARTICLES

De Broglie’s prediction was verified experimentally by C.J. Davis-


son and L.H. Germer and also independently by G.P. Thomson in
1927. To understand the Davisson-Germer experiment (see Fig. 2.1,
let us first consider a similar experiment using light waves. Clearly,
the classical experiments that we used to demonstrate the wave prop-
erties of light–Young’s double slit experiment or diffraction gratings–
will be difficult for such short wavelengths. For instance, on a good
quality diffraction grating with 10 thousand rulings/cm, the first or-
der maximum for a wave of wavelength λ = 100 pm the occurs at
θ = sin−1 λ/d = sin−1 10−4 ' 0.0057◦ , which is too close to the cen-
tral maximum. Clearly, a grating with d ≈ λ would do, which would
1168 CHAPTE R 38 PHOTONS AN D MATTE R WAVES
mean rulings at atomic distances apart.

electrons were sen


The apparatus was
Circular strate optical interfe
diffraction
Incident beam ring
to an old-fashione
(x rays or electrons) screen, it caused a fl
The first sever
Target interesting and se
(aluminum
crystals)
However, after man
Photographic apparatus, a patte
film where many electr
(a) had hit the screen. T
wave interference.
Figure 2.1: tus as a matter wa
eled through one
through the other s
ability that the elec
2.1.1 X-ray diffraction screen, hitting the s
The typical energy of photons of wavelength around λ = 100 pm gions
is correspondin
E = hc/λ ' 10 keV. An electromagnetic wave with such short wave- few electrons ma
fringes.
length and high energy is called X-ray. In 1912, Max von Laue arrived
at the idea of employing, as a diffraction grating for X-rays, a solidSimilar interfe
neutrons, and vari
body with regularly-arranged molecules, e.g. a crystal. The antici-
iodine molecules I
pated diffraction pattern by von Laue was indeed found by his exper-
sive than electrons
imenter colleagues W. Friedrich and P. Knipping. They admitted a
strated with the ev
thin beam of X-rays into a carefully oriented lead box surrounded by
and C70. (Fulleren
photographic films behind the box and on the sides. The intensity
arranged in a struc
maxima which had been anticipated by von Laue became evident in in
C60 and 70 carbo
the form of blackened spots on the film positioned behind the crystal.
as electrons, proton
(b)
However, as we co
must come a point
ing the wave natur
maximum occurs at“diffraction grating” for x rays. The idea is t
dimensional
sodium chloride (NaCl), a basic unit of atoms (called the
throughout the array. Figure ml 36-28a "1 represents
(1)(0.1 a sectio
nm)
NaCl and u ! sin"1
identifies sinThe unit cell is a
this basic!unit.
36-7 each
X-RAYside. DI FFRACTION d 1105 3000 nm
CHAPTER 2. PARTICLES When an x-ray beam enters a crystal 33 such as NaCl, x r
This is too close to—the
is, redirected central
in all maximum
directions to structure.
by the crystal be prac
gure 36-27 shows that x rays are produced when desirable,scattered
but, because x-raydestructive
waves undergo
C wavelengths are abou
interference, resulti
in other indirections
forthe interference is constructive, resulti
filament F are Max
accelerated
von Laueby was
a potential
awarded such
differ-
the gratings
Nobel prize cannot
1914 be constructed
“his discoverymechanically.
This process of scattering and interference is a form of dif
. In 1912,Fictional
of the diffraction of X-rays by crystals”. it occurred
Planes. to German
Although physicist
the process Max
of diffraction
ion grating cannot Thebetheory
used oftodiffraction
discriminate T to be in directions
solid, which consists of a regular array of atoms,
of X-rays complicated, the maxima
F turn out
in the x-ray by wavelength range. For
atomic crystals l ! 1 dimensional
is complicatedÅ be- “diffraction grating” for x rays. The ide
r example, Eq.cause36-25 when
shows an thatX-ray
the first-order
beam sodium
enters achloride (NaCl), a basic Na+ Cl–
unit of atoms (calle
Inciden
x ray
crystal such as NaCl shown inthroughout Fig. 2.2, the array. Figure X rays36-28a Wa0 represents a
the X-rays are scattered by the NaCl atoms
and identifies this basic unit. The unit cell
(1)(0.1 of
nm) the crystal in all directions. Yet we V d
! sin "1
! 0.0019#. each side.
3000 nm
can observe a diffraction pattern of Figure
al- 36-27 X rays are generated when such asd Na
When an x-ray beam enters aa0crystal
ternating intensity maxima
aximum to be practical. A grating with d ! l is, and electrons
min- (a) leaving heated
is redirected — in all directions byFthe filament are (b)
crystal struc
ima. The maxima appear in accelerated through a potential difference
directions
velengths are about equal to atomic diameters, scattered waves undergo destructive interference, r
as if the X-rays were reflected by planes V and strike a metal target T. The “win-
ed mechanically. in other directions the interference is constructive, r
in thevon
crystal–called dow” W in Ray the2Figure
evacuated 2.2:chamber C is
man physicist Max Laue that acrystal planes–that
crystallineThis process
transparent of scattering
Ray 1
to x rays. and interference is a form
r array of atoms,contain regular
might form arrays of the
a natural atoms such
three- Fictional Planes. Although θ θ the process of diffr
” for x rays. Theasidea
shown in Fig.
is that, in a2.3, left. such
crystal Theseas planes are separated byθthe
complicated, the maxima turn out to be in direc
θ
distance
d that the
unit of atoms (called happens to be
unit cell) equalitself
repeats to the size of thed unit cell dimension a0 .
sin θ d sin θ
-28a represents a section
Rays through
i is reflected from the ithofplane (i = 1, 2 or 3).d The
a crystal angles of in- d
θ θ
unit. The unit cidence
cell is aand
cubethemeasuring a0 on by θ are measured from the
reflection, denoted Na+planes Cl– d In

(and not from the normal as we did in geometrical optics).


The extra distance ofLooking
ray 2
crystal such asatNaCl,
the xdrawing
rays areon scattered
the right of Fig. 2.3(c)we seedetermines
— that that the the path
interference.
differ-
a0 (d)

by the crystal structure.


ence betweenIn some directions
the rays the from two
reflected adjacent planes is 2d sin θ,
Figure 36-28 (a) The cubic structure of NaCl, showing the sodium
tive interference, resulting
hence we expectin intensity
intensityminima;
maxima in thea unit directions determined
cell (shaded). (b) Incident x rays by undergo
the diffraction by th
ce is constructive, resulting in intensity maxima.
condition x rays are diffracted as if they were reflected by a family of paral
measured relative to the planes (not a0relative to a normal as in op
erference is a form of diffraction. 2d sin θ = mλ .
difference between waves effectively (2.1.1)
reflected by two adjacent p
the process of diffraction of x rays by a crystal(a) is A different orientation of the incident x rays relative (b) to the struc
of parallel planes now effectively reflects the x rays.
out to be in directions as if the x rays were
3 2 1
Cl– Incident
x rays Ray 2 Ray 1

a0 θ θ
θ θ
d θ θ θ θ
d
0 d sin θ d sin θ
d θ θ
(b) θ θ

The extra distance of ray 2


(c)
Figure 2.3: determines the interference. (d)

Figure 36-28 (a) The cubic structure of NaCl, showing the so


Eq. (2.1.1) is called Bragg’s law, and can serve as a powerful tool
a unit cell (shaded). (b) Incident x rays undergo diffraction
x rays are diffracted as if they were reflected by a family of
d
measured relative to the planes (not relative to a normal a
d θ
difference between waves effectively reflected by two adja
complicated, the maxima turn out to be in directions as if the x rays were
3 2 1
Na+ Cl– Incident
x rays
34 CHAPTER 2.
θ
PARTICLES
θ
a0

d θ θ
for either measuring the wavelength of X-rays, if the interplanar dis-
tance is known, or more a0 importantly, to study the atomic structure
d θ θ
(a) of crystals. It was derived by W.L. (b)Bragg, who shared with his father
W.H. Bragg the Nobel prize in 1915 for their “services in the analysis
of crystal structure by means of X-rays”.
The actual
Ray 2
process is differ-
Ray 1
ent from the simplified picture
of reflection θ by θ planes, which
nevertheless gives θ θ the correct re-
sult.
d Even this simplified picture d
predicts d asin θcomplicated
d sin θ
diffrac- d θ
tion pattern. θ AsθFig. 2.4 shows, d θ θ
there are other crystal planes in θ θ
The extra distance of ray 2 θ
thedetermines
same crystal. Although the
the interference.
(c) (d)
crystal is oriented the same way,
these
Figure 36-28 crystal
(a) The planes have
cubic structure differ-
of NaCl, showing the sodium and chlorine ions and
ent interplanar distance d
a unit cell (shaded). (b) Incident x rays undergo anddiffraction by the structure of (a). The
x rays are diffracted
are orientedas ifdifferently.
they were reflected by a family of parallel planes, 2.4:
Figure with angles
measured relative to the planes (not relative to a normal as in optics). (c) The path length
difference between waves effectively reflected by two adjacent planes is 2d sin u. (d)
2.1.2
A different Diffraction
orientation of the incident of elec-
x rays relative to the structure. A different family
of parallel
tronplanesbeam
now effectively reflects the x rays.
on crystals
The Davisson-Germer experiment is an exact analogue of the X-ray
diffraction on crystals, with the only difference that the incident beam
consists of electrons. For target Thomson used metal film while Davis-
son used crystal grid. They observed the same diffraction pattern as
can be observed using X-rays, which proved de Broglie’s hypothesis
of matter waves. Thomson and Davisson shared the Nobel prize in
1937 for their “experimental discovery of the diffraction of electrons
by crystals”.
Assuming that electrons can be considered matter waves and us-
ing de Broglie’s formula, the wavelength can be controlled well by
using definite acceleration voltage. Thus one can make a precise mea-
surement of the interplanar distances in the crystal by measuring the
positions of the maxima at several different electron energies.

2.1.3 Double-slit experiment with electrons


The double-slit experiment was performed with a beam of electrons in
1961 by C. Jönsson resulting in the expected interference. Later the
CHAPTER 2. PARTICLES 35

experiment was repeated such that there was only at most one electron
at any time between the source and the detector screen by P.G. Merli
and his collaborators in the 1970s and later by A. Tonomura and his
collaborators in 1989 (see [Link] oWRI-
LwyC4). The technical details are not in our interest here. The im-
portant message is that even single electrons pass through the two
slits such that if many electrons do so independently, nevertheless an
interference pattern appears. So clearly, not the ensemble of elec-
trons behave as classical waves, but each electron separately behaves
as some sort of a wave. We call those waves matter waves.
We usually think of particles as little projectiles that cannot be
divided further. Classically, such projectiles path through one or the
other slit. Let us denote with Pi (x) the probability that a particle
passes through slit i (i = 1, 2) such that the other slit is closed and
reaches the detector at a distance x from the central point on the
screen that lies on the line connecting the source and the slits. Then
the probability distribution of the position of arrival on the screen
when both slits are open is

P12 (x) = P1 (x) + P2 (x) , (2.1.2)

so the probabilities sum up without interference. If we used waves


instead of particles, we would know that the intensities Ii (x) of the
waves corresponding to the cases when only slit i is open, sum up with
interference when both slits are open, i.e.
q
I12 (x) = I1 (x) + I2 (x) + 2 I( x)I2 (x) cos δ

where δ is the phase difference between the spherical wavelets coming


from slits 1 and 2 to the screen. The double-slit experiment with
electrons demonstrates that P12 (x) 6= P1 (x) + P2 (x) for elementary
particles (but such experiments were even performed also with large
molecules consisting 810 atoms).
In 2012 R. Bach and his collaborators performed the double slit
experiment with electrons such that using a movable mask they were
able to control which slit the electron passed through. The simplified
set-up is shown in Fig. 2.5. (a) An electron beam passes through a
wall with two slits in it. A movable mask is positioned to block the
electrons, only allowing the ones traversing through slit 1 (P1 ), slit 2
(P2 ), or both (P12 ) to reach the backstop and detector. (b,c) Proba-
bility distributions are shown, (Experimental in false-colour intensity)
36 CHAPTER 2. PARTICLES
Controlled double-slit electron diffraction 6

Figure 1. Simplified setup. a, An electron beam passes through a wall with two
slits in it. A movable mask is positioned to block the electrons, only allowing the
ones traversing through Figure
slit 1 (P1 ),2.5:
slit 2 (P2 ), or both (P12 ) to reach the backstop
and detector. b,c, Probability distributions are shown, (Experimental in false-colour
intensity) for electrons that pass through a single slit (b), or the double-slit (c). Inset
1,2, Electron micrographs of the double-slit and mask are shown. The individual slits
for electronsare
that50 nmpass
wide ×through a single
4 µm tall with a 150 nmslit (b),structure
support or themidway
double-slit (c).
along it’s height,
and separated by 280 nm. The mask is 5 µm wide × 20 µm tall. Reprinted from The
Feynman Lectures on Physics, Volume III, by Richard P. Feyman, Robert B. Leighton,
The detailed results
and Matthew of Available
Sands. Bach’sfrom experiment
Basic Books, an areimprint
shown inPerseus
of The Fig. 2.6.
Books
Group. Copyright ⃝ c 2011
On the left we see that the resulting probability distributions on the
screen as the mask is moved over the double-slit. The mask allows
the blocking of one slit, both slits, or neither slit in a non destructive
way. On the right we see how the electron interference builds up
from individual electrons. The bright spots indicate the locations of
detected electrons. Shown are intermediate build-up patterns from
the central five orders of the diffraction pattern (P12 ) with 2, 7, 209,
1004, and 6235 electrons (a-e).
Let us now imagine the following thought experiment: instead
of covering one of the slits, we put a source of intense light behind
the slits. Electric charges are known to scatter light, so when an
electron passes through one of the slits, then we can observe a flash
of light coming from behind that slit. In this experiment whenever an
electron reaches the screen, we also obeserve a flash from one of the
slits, so we conclude that the electron passes through only one slit at
a time. However, if we observe the probability pattern on the screen
in this experiment, we find that the interference disappears and the
probabilities simply add as in Eq. (2.1.2). If we now switch off the
intense light behind the slits, the interference patterm on the screen
reappears.
The message of the thought experiment is that the observation of
CHAPTER 2. PARTICLES 37

Controlled double-slit electron diffraction

Controlled double-slit electron diffraction 7

2760 nm

2640 nm

2540 nm
P1

2500 nm

2240 nm

1620 nm
P12

−140 nm

−1440 nm

−2320 nm
P2

−2500 nm

−2540 nm

−2580 nm

−2680 nm

Figure 2. Mask movement. A mask is moved overFigure 3. Buildup


a double-slit (inset)ofand
electron
the diffraction. “Blobs” indicate the locations of detecte
resulting probability distributions are shown. The mask electrons.
allows theShown are intermediate
blocking of one build-up patterns from the central five orders
theThe
diffraction pattern (P1250) magnified from figure 2, with 2, 7, 209, 1004, and 623
Figure 2.6:
slit, both slits, or neither slit in a non destructive way.
electrons
individual slits are
(a-e). A The
full labeled
movie of the electron build-up is included in the supplementar
nm wide and separated by 280 nm. The mask has a 5 µm wide opening.
dimensions are the positions of the center of the mask. dataP(see Supplementary Movie 2)
1 , P2 , and P12 are the
probability distributions shown in figure 1. (See Supplementary Movie 1 for more
positions of the mask.)
38 CHAPTER 2. PARTICLES

electrons changes their path. This should not be a big surprise as light
is electromagnetic wave and the electromagnetic field exerts a force
on the electrons. How can we reduce the effect of light on the motion
of the electrons? We already know that light consists of photons, with
energy and momentum proportional to the frequency of light. We can
reduce the effect of light on the electrons if we reduce its momentum,
i.e. its frequency, or increase its wavelength. Indeed, increasing the
wavelength of the light beyond a certain value the interference pattern
on the screen reappears. This value is comparable to the distance
between the slits, but this light cannot resolve the position of the
slits, so we cannot tell any longer where the electron passed through.
We can conclude from the real and
imaginary experiments that we can-
not observe the path of the electrons
without destroying the interference on
the screen. This non-trivial statement
was formulated by W. Heisenberg into
the principle of uncertainty. Accord-
ing to this principle, the laws of Nature
would be contradictory unless there is
a fundamental limitation in the preci-
sion of our measurements. In the case
of double-slit experiment with electrons
this means that it is not possible to de-
vise an experiment such that we can tell
which slit the electron passes through Figure 2.7:
without destroying the interference on
the screen. The original problem that lead Heisenberg to his con-
clusion was that the path of an electron can be followed in a cloud
chamber (see for instance the trace of the first ever observed positron
in Fig. 2.7), so it seems that the notion of classical path can be em-
ployed. However, the notion of the classical path of an electron in-
side an atom has disastrous consequences as that makes the atom
unstable (we shall give more details in the next section). In order
to resolve this paradox, Heisenberg suggested to use only observable
physical quantities to descibe the properties of the electrons inside the
atom. The classical path (i.e. position and velocity) is not so. It is
not observable because the uncertainties in the simultaneous measure-
ment of position and momentum cannot be made arbitrary small, but
∆x∆px ≥ ~ = h/2π. If we want to localize the electron to a fraction
CHAPTER 2. PARTICLES 39

of the size of the atom, then its momentum becomes so large that it
leaves the atom immediately, i.e. the atom disappears similarly as the
interference pattern when we observe which slit the electron passes
through.
The modern viewpoint is that both matter and radiation are phys-
ical fields such as the electromagnetic field, which takes into account
both the relativistic and the quantum effects. These fields can show
both wave and particle properties. The latter are called elementary
excitations of the field or particles for short. In these lectures we
shall concentrate only on the experiments that lead to the develop-
ment of non-relativistic quantum mechanics–the understanding of the
quantum effects. For that purpose first we explore the structure of
atoms.

Exercises
Matter waves

1. X-rays of wavelength λ = 120 pm undergo second-order reflec-


tion at a Bragg angle of θ = 28.1◦ from a LiF crystal. What is
the interplanar distance of the reflecting planes in the crystal?
2. Consider a two-dimensional square crystal structure, such as one
side of the structure of NaCl as shown in the lecture notes. The
largest interplanar spacing of reflecting planes is the unit cell
size a0 . Compute
(i) the second largest;
(ii) the third largest;
(iii) the fourth largest;
(iv) the fifth largest interplanar spacing. Do you recognize any
pattern?
3. Compute the de Broglie wavelength of an electron accelrated in
the Large Electron-Positron collider to a total energy of E =
45.6 GeV.
4. We masure the interplanar distances in a certain crystal using
electron diffraction. We find that the radii of the first-order cir-
cular maxima (on the screen at a distance of ` = 13.5 cm from
the target) for two differently oriented crystal planes, R1 and
R2 depend on the electron energy as shown in Table 2.1. Deter-
mine the interplanar distances and their statistical uncertainties
belonging to R1 and R2 .
40 CHAPTER 2. PARTICLES

Table 2.1:
U/kV R1 /cm R2 /cm
3.0 1.46 2.55
3.5 1.38 2.48
4.0 1.26 2.21
4.5 1.15 2.13
5.0 1.13 1.93

5. A beam of atoms emerges from an oven of temperature T =


400◦ C. The distribution of the speeds of the atoms follows the
Maxwell-distribution. Compute the (i) average and (ii) most
probable de Broglie wavelength of the atoms.
Chapter 3

Atoms

In the course on Thermodynamics based on the atomic model of Dal-


ton we developed a model of matter: the kinetic model that was
very successful in explaining macrospcopic phenomena (such as sur-
face tension, mixing and with some refinement thermal equilibration)
and properties of matter (such as the dependence of internal energy
on temperature) in a quantitative way. In these models the basic con-
stiuents were assumed to be structureless and indivisible. With the
discovery of the electron by J.J. Thomson in 1897 (he was awarded the
Nobel prize in 1906), which can emerge from atoms, it became clear
that this simplistic view of atoms was not tenable any longer and spec-
ulations started about what and how builds the atoms. Physicists ar-
gued that atoms must also contain some positive charge because they
were electrically neutral, but nobody knew how this positive charge
was distributed inside the atom. An historic breakthrough in this en-
devour was made by Ernest Rutherford and his two young collegues,
Hans Geiger and Ernest Mardsen in 1911, which became known as
Rutherford’s experiment. At the time Rutherford was already a No-
bel laurate as he was awarded the chemistry prize in 1908.

3.1 Rutherford’s experiment


3.1.1 Setup and result

41
42 CHAPTER 3. ATOMS
Atomic nucleus and atom are not the same1234. An atom is the smallest unit of matter that has the properties of an element3. An atom
42-1
consists of two regions: the nucleus and the electron cloud24. DISCOVE
The nucleus RofI NG
is the center TH
the atom Econtains
and N UCLE US
protons 127
and neutrons
The actual experimental setup is shown
ew) used in Rutherford’son the leftlaboratory figure of Fig. in 1911 3.1.– 1913 In 1911 to it
was already known
by thin metal foils. The detector can be rotated to vari- that some elements,
Alpha source
f. The alpha source called wasradioactive,
radon gas, a decay produce product radiation.
of
p” apparatus, theThree atomic types nucleus of radiation
was discovered. were identified:
α, β and γ. The α rays were known to
have relatively large mass (about 7300
times more massive than the electron)
and positive electric charge (−2 times Gold foil
s known thatthe certain
charge elements,
of the electron). called radioactive, Today we φ
spontaneously,know emitting thatparticlesthese are in the the process.
nuclei ofOne the
mits alpha (a) helium particles [Link] have Usingana energy source of of α aboutrays
42-1 DISCOVE R I NG TH E N UCLE US 1277

ese particles arefrom helium decaying nuclei. radon enclosed in a glass


Figure 42-1 An arrangement (top view) used in Rutherford’s laboratory in 1911 – 1913 to
Detector
direct energetic tube,
study the scattering alpha
of a Geiger
particlesparticles
and
by thin atThe
metal Mardsen
foils. a detector
thinbombarded
target
can be rotatedfoil
ous values of the scattering angle f. The alpha source was radon gas, a decay product of
to vari- Alpha source

With athis thin


simple foil
ch they were deflected as they passed through the
radium. “tabletop” of gold
apparatus, the and
atomic measured
nucleus was discovered. the
e about 7300distribution times moreofmassive deflectedthan α particles
electrons, as
a function of the scattering angle φ. Gold foil
In Rutherford’s
They day were it wasastonished
known that certainto elements,
explore called radioactive,
that Figureφ 3.1:
perimental
transformarrangement of Geiger
into other elements spontaneously, emittingand particles Marsden.
in the process. One
Number of alpha particles detected

such element sometimes


is radon, which emits the alphaα(a)particles
particles that have were scat-
an energy of about
n-walled glass now
5.5 [Link] tube of these
know that radon particlesgas.
are heliumThe experiment
nuclei. Detector
tered backwards.
of alpha particles
and measure the The
thatwhich
extent tofull
areresultdeflected
they were deflected
through
Rutherford’s idea was to direct energetic alpha particles at a thin target foil
as they
of the distribution passed
vari-
through the 107 Theory
foil. Alpha particles, which are about 7300 times more massive than electrons,
have a charge ofof scattered
!2e. particles is shown on the 10 6 Data
esults. Note especially
Figure 42-1 shows thethat the
experimental
right if Fig. 3.2. This result was sur- vertical
arrangement scale
of Geiger is andlog-
Marsden.
Number of alpha particles detected

Their alpha source was a thin-walled glass tube of radon gas. The experiment
f the particles arethescattered through rather small 105
prising
involves counting because
number of alpha the
particlesprevailing
that are view
deflected through atvari- 10 Theory
7

ous scattering angles f.


e big surprise — shows
that
Figure 42-2 a time
very small
was
their results. Notefraction
the modelthat
especially ofof them
J.J.
the Thom-
vertical are
scale is log-
10 104Data 6

5
ngles, approaching son. 180°. He advocated In Rutherford’s
arithmic. We see that most of the particles are scattered through rather small
that the
angles, but — and this was the big surprise — a very small fraction of them are
words:
positive “It 10
10 10
3 4

event that
scattered evercharge
through happened inside
very large angles,to theme atom
approaching in180°.
my was
In life. It was
homoge-
Rutherford’s words: “It 10
2
3
was quite the most incredible event that ever happened to me in my life. It was
10 10
had firedalmost aas15-inch
neously
incredible asshell you hadatfired
if distributed a apiece through
15-inch of attissue
shell theofpaper
a piece entire
tissue paper
2

and it [the shell] came back and hit you.” 10


d hit you.” volume of the atom, and the electrons
Why was Rutherford so surprised? At the time of these experiments, most
0° 10
20° 40° 60° 80° 100°120°140°
0°Scattering
20° angle
40°φ 60° 80° 100°120°140°
urprised? At were the time
physicists believed thought
in the
of these
so-called to vibrate
plum
experiments,
pudding at fixed points.
model of
been advanced by J. J. Thomson. In this view the positive charge of the atom was
the atom,
most
which had
Scattering angle φ
Figure 42-2 The dots are alpha-particle

alled plum to The


thoughtpudding be spread kinetic
outmodel
throughenergy
of
the the
entireofvolume
the ofαtheparticles
atom, which
atom. Thehad in
electrons scattering data for a gold foil, obtained by
Geiger and Marsden using the apparatus of
(the “plums”) were thought to vibrate about fixed points within this sphere of
n. In this view the(theexperiment
chargethe positive charge was Ekof'the 5.5 atom
MeV. was Ac- Figure 42-2solidThe
Fig. [Link] curve isdots are alpha-particle
the theoretical
prediction,Figure
based on the 3.2:
positive “pudding”). assumption that the
ugh thethrough cording
The maximum
entire to
deflecting
volume Rutherford’s
force
of the
that could act on an
atom. estimate
The
alpha particlein as itor-
passed scattering data for a gold
chargedfoil, obtained by
far tooelectrons
atom has a small, massive, positively
such a large positive sphere of charge would be small to deflect the [Link] data have been adjusted to fit the
der to asfind α1°.particles scattered back- Geiger andat theMarsden
experimental using the apparatus o
o vibrate about fixed points within this sphere of
alpha particle by even much as (The expected deflection has been compared theoretical curve point
to what you would observe if you fired a bullet through a sack of snowballs.) The
wards the positive charge inside the atom Fig.
[Link]
that is enclosed
be solid
in a circle. curve
concentrated is the theoretical
”). electrons in the atom would also have very little effect on the massive, energetic
prediction, based of on such
the assumption that th
alpha [Link]
They would, a sphere of small
in fact, be themselves radius
strongly [Link]
deflected, as a potential energy a
orce thatswarm could act on
of gnats would
homogeneously an alpha
be brushed aside by a particle as it them.
stone thrown through passed atom has
theα aparticles
small, massive, positively charge
Rutherford saw that, to deflect thechargedalpha particlesphere as
backward, there amustfunction
be a ofIncident distance r from
phere oflargecharge
force; thiswould
theforce centercould bebeoffarthetoo
provided small
if the
sphere istocharge,
positive deflectinsteadthe of being [Link] Target data have been adjusted to fit
spread throughout the atom, were concentrated tightly at its center. Then the
as 1°. (The
incoming expected
alpha particle could deflection
get very closehas to thebeen compared
positive charge without pene- theoretical
foil
curve at the experimental point
2 2
 
you fired aFigure bullet through a sack of snowballs.)
trating it; such a close encounter would result in a large deflecting 1 2Ze
force. 1
The r
that is enclosed in a circle.
42-3 shows possible paths taken Ebyp (r) =alpha particles as they pass3 − 2 .
typical (3.1.1)
lso havethrough
very littleof the
the atoms effect onAs the
target foil. we see,massive, 4πεenergetic
most are either 0 R or
undeflected 2 only R
slightly deflected, but a few (those whose incoming paths pass, by chance, very close
fact, be
to athemselves strongly
nucleus) are deflected through large deflected,
angles. From anmuch
analysis ofas
the a
data,
Rutherford concluded that the radius of the nucleus must be smaller than the radius
ed aside byatom
of an a by
stone
a factorthrown through
of about 10 . In other words, them.
4
the atom is mostly empty space.
CHAPTER 3. ATOMS 43

In order that the α particle be backscattered, the maximum of this


potential energy
3Ze2
Emax = Ep (0) =
4πε0 R
has to be larger than the kinetic energy of the α particle, Emax > Ek .
Thus the positively charge matter has to be confined into a sphere of

3Ze2
R< ' 71.5 fm ,
4πε0 Ek
which is much smaller than the radius of the gold atom (about
129 pm). Thus Rutherford decided to derive the differential cross
section of a particle of charge ze on another particle of charge Ze
(e being the unit charge and z, Z are positive integers), which we also
do next.
The repulsive force force between the two positively charged point-
like particles is given by Coulomb’s law,

1 zZe2
F = .
4πε0 r2
This is a conservative force so there exists a corresponding potential
energy that is inversely proportional to the distance r between the
two particles, Ep ∝ 1/r. In the case of this potential energy, the path
of the scattered particle is a hyperbola, which we prove first.

3.1.2 Algebraic properties of a hyperbola


Let us first collect some algebraic properties of
the hyperbola depicted in Fig. 3.3. The hyper-
bola consists of two infinite lines that can be
positioned symmetrically with respect to the x
and y axes in the x − y plane such that their
closest points V1 and V2 are at a distance 2a
apart. These points lie on the major axis that
we choose to coincide with the x axis. At half
way between these points is the center M of the
hyperbola that we choose to lie at the origin of
our coordinate system. At distance c > a from
the center there are the two foci F1 and F2 .
With this positioning the points P of the
hyperbola in the x − y plane are given by the
equation
Figure 3.3:
x2 y2
− 2 =1 (a2 + b2 = c2 ) , (3.1.2)
a2 b
44 CHAPTER 3. ATOMS

or
bp 2
y(x) = ± x − a2 (3.1.3)
a
where the positive solution corresponds to the line above and the negative one to
that below the x axis. The hyperbola is symmetric with respect to the x axis and
crosses it at |x| = a where y = 0. The foci lie in points (±c, 0) where

b2
 
c2 = a2 + b2 = a2 1 + 2 ≡ a2 ε2 . (3.1.4)
a
p
The parameter ε = 1 + b2 /a2 > 1 characterizes the eccentricity of the hyper-
bola, the smaller ε, the closer its asymptotes (lines y = ± ab x) to the x axis.
The parametric equations of the hyperbola are

x = a cosh u , y = b sinh u (3.1.5)

where the parameter u has meaning related to area. Let us consider the part of
the hyperbola in the upper half plane where y > 0. The position vector to the
point (x, y(x)) and the [0, x] section form two sides of a rectangular triangle whose
area is t = xy(x)/2. This triangle is cut into two parts by the hyperbola. Let us
denote the area of the part below the hyperbola by t1 , while that of the other part
by t2 , so t = t1 + t2 . The area t1 is simply
Z x p Z x/a p
b
t1 = x2 − a2 dx = ab x2 − 1dx
a a 1
  
ab x y(x) x y(x)
= − ln +
2 a b a b

xy(x) ab
= − ln(cosh u + sinh u)
2 2
ab
=t− u. thay vi chon vi tri tu goc toa do den P thi
2 thuong trong bai toan vat chuyen dong trong
Then truong luc, ta chon vecto vi tri tu vat 2 (>> vat
ab 1) den P
t2 = t − t1 = u , (3.1.6)
2
so u is the measuring number of the area t2 in units of 21 ab.
The radius vector is the vector from the more distant focus to the hyperbola.
The area of the triangle enclosed by the radius-vector, the x axis and the position
vector is
1 1
t0 = cy(x) = aε b sinh u . (3.1.7)
2 2
Finally, size of the area enclosed by the radius-vector, the x axis and the hyperbola
is
ab
t0 + t2 = f (3.1.8)
2
where
f = u + ε sinh u (3.1.9)
1
is the measuring number of this area in units of 2
ab.
CHAPTER 3. ATOMS 45

3.1.3 Motion in an 1/r repulsive potential


Let us assume that a particle of mass m and charge ze approaches
from infinity a target of infinite mass and charge Ze, placed at the
origin. The force is a central force, so the magnitude of the angular
momentum L ~ of the projectile is a constant of the motion. The po-
sition vector is perpendicular to the angular momentum, so the path
lies in the plane orthogonal to L.~ Let this plane be the x − y plane.
The velocity ~v lies also in this plane.
We also assume that the trajectory of the projectile is on the half
of a full hyperbola that lies fully in the x > 0 half plane. With this
choice of the coordinate system, the position vector coincides with the
radius vector of Sect. (3.1.2)

~r = (c + a cosh u) ~i + y ~j = a(ε + cosh u) ~i + b sinh u ~j . (3.1.10)

The length of the position vector is computed from the components


as

r2 = a2 (ε + cosh u)2 + b2 sinh2 u

= a2 (ε2 + 2ε cosh u + cosh2 u) + a2 (ε2 − 1)(cosh2 u − 1)


(3.1.11)
= a2 (1 + 2ε cosh u + ε2 cosh2 u)

= [a(1 + ε cosh u)]2 ,

so
r = a(1 + ε cosh u) . (3.1.12)

Clearly, the limits u → ∞ and r → ∞ are equivalent.


The velocity is the derivative of the position with respect to time,
so
˙ = dr du = u̇(a sinh u ~i + b cosh u ~j) .
~v =~r (3.1.13)
du dt
The speed squared can be computed from the velocity components as

v 2 = u̇2 [a2 sinh2 u + b2 cosh2 u]

= u̇2 [a2 (cosh2 u − 1) + a2 (ε2 − 1) cosh2 u] (3.1.14)

= u̇2 a2 (ε2 cosh2 u − 1) .


46 CHAPTER 3. ATOMS

The magnitude of the specific angular momentum,


~
|L|
l= = |~r × ~v | = u̇ab(ε cosh u + cosh2 u − sinh2 u)
m (3.1.15)
= ab u̇(1 + ε cosh u)

is a constant of the motion. The conservation of l is equivalent to the


law of areas: the radius vector sweeps out equal areas in the plane of
the projectiles orbit in equal time intervals, or using Eq. (3.1.9)

f˙ = u̇ + εu̇ cosh u = ω (3.1.16)

is a constant of the motion. Expressing u̇ from Eq. (3.1.16) and sub-


stituting it into Eq. (3.1.14), we find
ε cosh u − 1
v 2 = ω 2 a2 . (3.1.17)
ε cosh u + 1
As
ε cosh u − 1
lim = 1, (3.1.18)
u→∞ ε cosh u + 1
the speed at infinity is v∞ = ωa. Using Eq. (3.1.16), we can rewrite
Eq. (3.1.15) as
l = ab ω , (3.1.19)
hence v∞ = l/b.
The total mechanical energy Et = Ek + Ep is also a constant of the
motion. Choosing the zero point of the potential energy at infinity,
we obtain this constant as
1 2 ml2
Et = 0 + mv∞ = 2 > 0. (3.1.20)
2 2b
Finally, we can express the potential energy Ep = Et − Ek as

1 l2 1
Ep = m 2 − mv 2
2 b 2
ml2 ε cosh u − 1 ml2

1
= 2 1− = 2 (3.1.21)
2b ε cosh u + 1 b ε cosh u + 1
ml2 a
= 2
b r
where we used Eq. (3.1.12). Eq. (3.1.21) is our result: in case of a
central force, the motion on a hyperbola implies that the potential
CHAPTER 3. ATOMS 47

Figure 3.4:

energy is inversely proportional to the distance r of the particle from


the force centre. At given initial conditions the potential energy de-
termines the trajectory uniquely, hence we can reverse our result: the
trajectory in a 1/r repulsive potential, with total mechanical energy
E > 0 is a half hyperbola as shown in Fig. 3.4.
Let us assume that the projectile approaches from the direction
(−∞, +∞) and leaves to the direction (+∞, +∞). If we denote the
angle between the asymptotes of the incoming and outgoing directions
and the x axis by ∓χ, then
b p
tan χ = = ε2 − 1 , (3.1.22)
a
or expressing the eccentricity,
ε = cos−1 χ . (3.1.23)
The distance b0 of the incoming asymptote from the centre of the force
is called impact parameter, and simple geometry yields
b0
sin χ = . (3.1.24)
c
The scattering angle is ϑ = π − 2χ, and we obtain readily the relation
between the impact parameter and the scattering angle:
1
b0 = c sin χ = aε sin χ = a tan χ = a = b. (3.1.25)
tan(ϑ/2)
Thus, the impact parameter equals b, so we do not have to use different
symbols for these.
48 CHAPTER 3. ATOMS

3.1.4 Differential cross section of scattering on a


point-like scattering centre
The number of projectile particles in a section of central angle dϕ
on a ring of radius b, width db is Φbdbdϕ where Φ is the flux of the
projectile particles (assumed to be homogeneous). As the scattering
is elastic, this number is equal to the number of scattered particles
into the solid angle dΩ = d cos θ dϕ,

Φ d cos θ dϕ = Φ bdb dϕ ,
dΩ
or after simplification
dσ db
=b
dΩ d cos θ
where dσ/dΩ is the differential scattering cross section. In Eq. (3.1.25)
we already found the relation between the impact parameter and the
scattering angle, so a simple differentiation immediately gives the gen-
eral result for the differential cross section in case of a central repulsive
force:
dσ ab db ab 1
=− = 2
dΩ sin ϑ dϑ 2 sin ϑ sin (ϑ/2)

a2 1
= (3.1.26)
4 sin(ϑ/2) cos(ϑ/2) tan(ϑ/2) sin2 (ϑ/2)

a2 1
= .
4 sin4 (ϑ/2)

We obtain the Rutherford scattering formula if we apply the gen-


eral formula in Eq. (3.1.26) to the case of Rutherford scattering. The
potential energy is that of the Coulomb force,

zZe2 ml2 a 2 a
Ep = K = 2 = mv∞ (K −1 = 4πε0 ) , (3.1.27)
r b r r
so
zZe2
a=K 2
. (3.1.28)
mv∞
The quantity 21 mv∞ 2
is simply the kinetic energy of the α particles
emerging from the source. The resulting cross section is depicted
as the solid line in Fig. 3.2. We see that the prediction agrees with
CHAPTER 3. ATOMS 49

Rutherford’s experimental results very well, which provided the ob-


servational basis for Rutherford’s model of the atom. According to
him the atom consists of a point-like (as compared to the size of the
atom), positively charged nucleus with electrons bound around it, like
a tiny solar system.

3.1.5 Differential cross section of scattering on an


extended scattering centre
In our derivation of the Rutherford scattering formula we used that
the target was point-like. More precise measurements revealed that at
large ϑ there are differences between the data and predictions. One
possibility to explain such differences is that the target is actually
not exactly point-like, but has a finite size. How does the finite size
modify the formula obtained for the point-like scattering centre?
We may represent the finite size of the nucleus with a charge dis-
tribution Ze%(~r) where the function %(~r) has to be normalized to one,
Z
%(~r)d3 r = 1 .

Here and below the integration extends over the whole space if it is
not stated otherwise explicitly. For a point-like target nucleus this
normalised charge distribution is a δ distribution, %(~r) = δ(~r), which
means that if ~r 6= 0, then % = 0.
The Fourier transformed of the normalized charge distribution,
Z
F (~q) = %(~r) ei~q·~r/~ d3 r

is called form factor. We state without proof that the differential cross
section on an extended target is related to that on a point-like target
according to the equation
   
dσ dσ
= |F (~q)|2
dΩ % dΩ R
where R refers to the Rutherford formula (scattering on point-like
target) and ~q = ~k − ~k 0 represents the momentum transfer between
the projectile and target during the scattering process. For point-like
target %(~r) = δ(~r) and
Z
F (~q) = δ(~r) ei~q·~r/~ d3 r = ei·0 = 1
50 CHAPTER 3. ATOMS

as required. Assuming spherically symmetric density distribution, %


depends only on the radius, %(~r) = %(r). In this case we can perform
the integrals over the polar and azimuthal angles,
Z ∞ Z 1 Z 2π
2 iqr cos θ/~
F (~q) = F (q) = r dr%(r) d cos θe dφ
0 −1 0

sin(qr/~)
Z
= r2 dr%(r) 2 2π .
0 qr/~
If the momentum transfer is small, we can truncate the Taylor expan-
sion of the sine at low orders. Keeping the first two terms only,
sin(qr/~) 1  qr 2
≈1+
qr/~ 6 ~
and the form factor is approximately
Z ∞
q2
Z 
2 4
F (q) ≈ 4π r dr %(r) + 2 r dr %(r)
0 6~
Z ∞
1 q 2 hr2 i
 
= 1− 4π r2 dr %(r)
6 ~2 0

1 q 2 hr2 i
=1−
6 ~2
R∞
because 4π 0 r2 dr %(r) = 1 according to the normalization condition.
The square root of the quantity
Z ∞  , Z ∞  Z
2 2 4
rrms ≡ hr i = %(r) r dr %(r) r2 dr = %(r) r2 d3 r
0 0

is called root mean squared radius of the nucleus. Our calculation


shows that for small rrms the differential cross section differs from
the Rutherford formula only for large momentum transfer, i.e. for
large scattering angles. This suggests that although the nucleus is
not point-like, it is very small. Its radius is in the femtometer range.
The ratio of the measured differential cross section to the Ruther-
ford formula provides experimental information on the form factor.
Then in principle, an inverse Fourier transformation reveals the charge
distribution. However, the inverse Fourier transformation is an inte-
gral over all momentum transfers, i.e. we need to know the cross sec-
tion at large scattering angles where it tends to vanish (see Fig. 3.2),
CHAPTER 3. ATOMS 51

hence difficult to measure precisely. In practise, we assume a func-


tional form of the charge distribution, described by a small number of
parameters, which are then fitted to the measured values of the form
factor.

3.1.6 Significance of the discovery of the nucleus


Rutherford’s experiment bears with tremendous significance for later
history of physics and of mankind.
First of all, this experiment was performed well before its time and
its success was partly due to pure luck. Above we derived the formula
for the differential cross section using the laws of classical physics, just
as Rutherford did himself. In 1911 he did not have other choice as the
laws of quantum mechanics were unknown. Today we know that for
a microscopic system classical mechanics is not applicable. Yet, the
computation of the differential scattering cross section of a charged
particle on a point-like charged target leads to the same formula both
in classical and quantum mechanics. Had it been not the case, the
interpretation of the measured cross section would have been much
more difficult and so would have been the acceptance of Rutherford’s
model of the atom.
As Heisenberg concluded in his uncertainty principle from trying
to understand the structure of the atom, the birth of quantum me-
chanics would have been likely delayed without having a firm (or so
thought) understanding of the atomic structure. Quantum mechanics
played essential role in the discovery and development of the physics
of nuclear fission reactions. Hence it is quite plausible that the devel-
opment of the atomic bomb during the second world war would not
have been possible if the classical and quantum predictions for the
Rutherford scattering formula had been different.
Secondly, Rutherford’s experiment was also a precursor of modern
experiments aiming at revealing the deep structure of matter. The
scattering method became an ubiquitous method in particle physics
and still today this is the key type of experiment in exploring the
microcosmos.

Exercises
Rutherford scattering
1. An α particle of kinetic energy Ek = 5.5 MeV collides with a gold
nucleus head-on. What will be the smallest distance between the
52 CHAPTER 3. ATOMS

two nuclei during this collision?


2. An α particle of kinetic energy Ek = 5.5 MeV collides with a
gold nucleus such that the incoming and outgoing asymptotes a
orthogonal. What will be the smallest distance between the two
nuclei during this collision?
3. An α particle of kinetic energy Ek = 5.5 MeV collides with a
gold nucleus such that the excentricity of its hyperbola path
is ε = 2. What will be the smallest distance between the two
nuclei during this collision?
4. Inorganic crystals, e.g. CsI(Tl), ZnS(Ag) are often used in thin
sheets as α particle monitors. We prepare a ring of such a sheet
with radius r = 1.0 m and width d = 4.9 cm. We have a beam of
protons from a cyclotron with kinetic energy Ek = 4.0 MeV and
intensity I = 40 µA, targeted on a thin (of thickness t = 20 µm)
aluminium foil (of density % ' 2.7 g/cm3 and molar mass M '
27 g/mol). The beam size on the target is A = 1 mm2 . How
many α particles will hit the ring detector in ∆t = 1 minute of
operation?
5. Imagine the gold nucleus as a sphere of radius R with uniform
charge distribution % = 8.8 µC
m3 .
(i) Compute the rrms radius of this nucleus.

3.2 Atomic spectra


We have already presented a detailed discussion of the spectrum of
the blackbody radiation, a famous puzzle at the end of the 19th cen-
tury. Although this spectrum is a continuous function of the frequency
of the emitted electromagnetic wave, its interpretation was given by
Einstein assuming that the radiation is emitted in the form of energy
quanta.
Around the same time
there was another puzzle
of electromagnetic spec-
tra that kept physicists
baffled for long. Already
in 1802 W.H. Wollaston
observed the appearance
Figure 3.5:
CHAPTER 3. ATOMS 53

of dark lines solar spec-


trum that were rediscovered by J. von Fraunhofer in 1814 (see
Fig. 3.5). In 1860 R.W.E. Bunsen and R.G. Kirchhoff found charac-
teristic emission lines in the emitted spectra of heated elements. An
example is shown in Fig. 3.6 where we see the emission lines of hydro-
gen in the visible spectrum. They also observed that the wavelengths
of the dark lines in the solar spectrum coincided with the those of
emission lines. Namely, the relations between the solar and emission
wavelengths are the following: λC = λred = 656.3 nm, λF = λblue =
486.1 nm, λf = λblue−violet = 434.1 nm, λh = λviolet = 410.2 nm.
Kirchhoff concluded that the dark lines in the solar spectrum were
caused by absorption of the solar radiation by chemical elements in
the solar atmosphere. This suggested that the elements could absorb
and emit electromagnetic radiation of specific wavelengths, character-
istic to the type of the element.
In 1885 J. Balmer found
an empirical law for the wave-
lengths of the visible emission
lines of hydrogen seen in Fig. 3.6:
n2
λ ' 364.6 nm . (3.2.1)
n2−4
In 1890 Rydberg gave a different
form of the same formula,
  Figure 3.6: Emission spectrum of
1 1 1
=R − 2 , (3.2.2) the hydrogen in the visible range
λ m2 n (on linear scale) and full range (on
logarithmic scale) of radiation.
with m = 2, n > m integer and
7
R ' 1.097 · 10 /m, which turned
out to catch the general feature of such series better. The series in
Eq. (3.2.2) is called the Balmer or Balmer-Rydberg series. Between
1906 and 1914 T. Lyman found hydrogen emission lines in the ul-
traviolet region (called Lyman series), which could be described with
the same formula (3.2.2), but with m = 1. In 1908 F. Paschen pub-
lished his finding of hydrogen spectral lines in the infrared region
(called Paschen series), which also obeyed the Balmer-Rydberg for-
mula, but with m = 3. The complete spectrum of hydrogen including
the Brackett series (discovered by F.S. Brackett in 1922 and corre-
sponding to m = 4) and Pfund series (discovered by A.H. Pfund in
1924 and corresponding to m = 5) and Humphreys series (discovered
54 CHAPTER 3. ATOMS

by C.J. Humphreys in 1953 and corresponding to m = 6) is shown in


Fig. 3.6 on a logarithmic scale of the wavelength.
It was also observed that for hydrogen-like ions with atomic num-
ber Z, but ionized Z − 1 times the same formula (3.2.2) remains valid
with the substitution R → RZ 2 .
Thus by the end of the first decade in the 20th century the va-
lidity of the empirical formula in Eq. (3.2.2) was established, but its
interpretation was a puzzle. In 1913 N. Bohr went to England to
work with J.J. Thomson and Rutherford. By that time the existence
of the electron and of the much more massive atomic nucleus was
known to them. It was natural to assume that the negatively charged
electrons were orbiting around the positively charged nucleus making
the atom neutral as seen from outside. However, such a (Rutherford)
model of the atom, envisaged like a miniature solar system bound by
the Coulomb force, cannot be realistic because classical accelerating
charges, like the electron on a closed orbit around the nucleus radiate
electromagnetic radiation.
According to the Larmor formula, the radiated power by an accel-
erating electric charge q is

1 q2 2
P = a
6πε0 c3

where a is the modulus of the acceleration. In the classical picture


of the Rutherford model this acceleration at a distance r from the
nucleus is given by Newton’s equation,

1 Ze2
= ma . (3.2.3)
4πε0 r2

So a2 ∝ r−4 vanishes quickly at large r, but becomes large at small


distances, which means that far from the nucleus the radiated power is
small and the orbit can be considered quasi-stable, meaning that the
radiated energy in one cycle is small compared to the total radiated
energy during the annihilation of the atom. For instance, an electron
orbiting around a proton at a distance of 100 pm looses relative energy
at a rate of about 10−6 per cycle. However, once the electron gets close
to the nucleus, it loses its energy much faster because the loss rate
per cycle is proportional to r−3 , so the atom is unstable. A numerical
estimate shows that it takes about 100 ps for a classical electron to fall
into a proton from a distance 100 pm, which shows that the Rutherford
CHAPTER 3. ATOMS 55

model of the atom cannot be correct because the hydrogen atom is


stable.
We can estimate the frequency of the radiation off an electron on
a quasi-stationary orbit easily because it is equal to the frequency of
the circular motion of the electron on this orbit, given by f = ω/2π
where ω is the angular speed of the electron. The latter is related to
the acceleration by the simple formula ω 2 = a/r, with r being the
radius of the orbit. Substituting the acceleration from Eq. (3.2.3) we
obtain the frequency as
r
1 ρ0 Z
f=
2π mr3
where we introduced the notation
e2
ρ0 =
4πε0
for this ubiquitous combination of constants. We can also express
this frequency as a function of the energy of the electron on this orbit
because
ρ0 Z
E = Ek + Ep = − , (3.2.4)
2 r
so s
1 8E 3
f= − 2 2 . (3.2.5)
2π ρ0 Z m
Lacking the explanation for the stability of the atom, and in order
to explain the Balmer-Rydberg formula, in particular to derive the
value of Rydberg’s constant R, Bohr made two postulates:
1. The postulate of stationary states: the hydrogen (and in general,
any) atom can exist for a long time without radiating in any one
of stationary states of well-defined energy.
2. The frequency postulate: the hydrogen (and in general any) atom
can emit or absorb radiation only when the atom changes from
one stationary state to another. The energy of the radiated or
absorbed photon is equal to the difference of these well-defined
energy levels,
hfmn = En − Em . (3.2.6)
Recall that Einstein’s assumptions in deriving the formula of
spectral-radiation were very similar in spirit to the second pos-
tulate.
56 CHAPTER 3. ATOMS

As stationary states are not allowed in a classical theory, and these are
related to the emission or absorption of photons, i.e. energy quanta,
we call them quantum states. The labels m and n are integers that
enumerate these quantum states, called quantum numbers. Multiply-
ing Eq. (3.2.2) with the universal constant hc = hf λ and comparing
the resulting equation to Eq. (3.2.6) we see that the energy levels
in the hydrogen atom are characterized uniquely by positive integer
principal quantum numbers n as
hcRZ 2
En = − (3.2.7)
n2
where we allowed for hydrogen-like ions with the inclusion of the factor
Z 2 , Ze being the charge of the nucleus.
In order to make connection between the electron moving on a
classical orbit far from the nucleus and the electron moving on the
stationary states inside the microscopic atom, Bohr also used the cor-
respondence principle: any generalization of a theory must agree with
an established theory of narrower validity in a well-defined limit. We
have already seen the application of this principle when we were look-
ing for a generalization of Galilean transformations to motions with
large speeds (as compared to the speed of light) such that the Lorentz
transformation formulae were to fall back to the Galilean ones in the
limit of small speeds. In the present case the correspondence principle
means that
the new (quantum) theory must agree with the classical
theory in the limit of large quantum numbers.
We now apply the correspondence principle to the electron jump-
ing between states with large quantum numbers, from state with quan-
tum number n to that with m = n−1 such that n → ∞. Bohr’s second
postulate in Eq. (3.2.6) together with Eq. (3.2.7) gives
2cRZ 2
 
2 1 1
fn−1,n = cRZ − ' ,
(n − 1)2 n2 n3
which should equal to the result of the classical computation for the
same frequency in Eq. (3.2.5), with energy taken from Eq. (3.2.7).
The resulting equation can be solved for Rydberg’s constant, yielding
a prediction for R in terms of fundamental constants:
me4
R= . (3.2.8)
8ε20 h3 c
CHAPTER 3. ATOMS 57

Substituting the known values for the fundamental constants, we ob-


tain numerically R = 10 973 731.568 508/m in excellent agreement
with the value obtained from spectroscopic measurements within the
uncertainty of the measurement.
The success of Bohr in explaining the emission spectra of elements
made physicists confident that Bohr’s postulates were correct. Never-
theless, this explanation was in a sense incomplete because the origin
of these postulates was not understood. Yet, it is interesting explore
its consequences that presumably should be taken seriously.

3.2.1 Quantized energy levels


Accepting that Rydberg’s constant is given by Eq. (3.2.8) exactly, we
can substitute it into Eq. (3.2.7) and find a purely quantum expression
for the energies of the stationary states of the electron in a hydrogen-
like ion,
ρ0 Z 2 1
En = − (3.2.9)
2 a0 n2
where for the sake of easing the formula we introduced the quantity
ε0 h2 ~2
 
h
a0 = = ' 52.92 pm ~= ,
πme2 ρ0 m 2π
called Bohr radius (having physical dimension of length). It would be
equal to the distance of the electron from the nucleus on the lowest
(n = 1) energy state in the classical theory (cf. the energy formula
with the classical one in Eq. (3.2.4)) if the correspondence principle
were valid for all quantum numbers, which is of course not true.
The first direct experimental support for the existence of quan-
tized energy states of electrons in the atom was provided in 1914 by
J. Franck and G. Hertz, who had not known about Bohr’s theory. The
experimental setup was simple, yet the results were so groundbreak-
ing that they were awarded the Nobel prize “for their discovery of the
laws governing the impact of an electron upon an atom”.
The experimental device (shown schematically and real in Fig. 3.7)
is a vacuum tube containing a mercury droplet. The tube is heated
(to a temperature 115◦ C) so the mercury evaporates into a vapor of
mercury atoms (of pressure about 100 Pa). There are three electrodes
in the tube: (i) a hot cathode C serves as an electron source, (ii) an
anode A collects the emitted electrons and between them there is a
(iii) metal mesh grid G, whose electric potential is positive to both
4
CHAPTER 3. ATOMS

the anode and the cathode, with UGC  UGA . In the experiment the
current through the anode is measured as a function of the voltage

Figure 3.7: Schematic and real pictures of the experimental apparatus

The result of our measurement performed in the class is shown in Fig. 3.8.
The current increased with increasing voltage up to about 4.9 V. This behavior is
typical of any vacuum tubes that don’t contain mercury vapor: the larger voltage
the larger anode current. At about 4.9 V the current dropped sharply, almost
vanished. The current then again increased steadily with increasing voltage until
about 9.8 V (= 4.9 V + 4.9 V) was reached where again a sharp drop was observed.
attached itself to the filament under highly repeatable This series of dips in current at about 4.9 V increments continued to voltage of
conditions.
This work was conducted during Advanced Lab
(PHY243W) at the University of Rochester.

Electronic address: [Link]@[Link]

Electronic address: [Link]@[Link]
[1] UR Advanced Laboratory Manual, The Franck-Hertz
Experiment [Online] [Link]
~AdvLab/2-Frank-Hertz/Lab02%[Link]
[2] P. Nicoletopoulos, Phys. Rev. E 78, 026403 (2008)
[3] R.E. Robson, [Link]., J. Phys. B 33, 507 (2000).
[4] G.F. Hanne, Am. J. Phys. 56 (8) (1988).
[5] D.R.A. McMahon, Am. J. Phys. 51 (12) (1983).
[6] E.B. Saloman, J. Phys. Chem. Ref. Data 35, 4 (2006)
[7] Yu. Ralchenko, Kramida, A.E., Reader, J., and
UGC between the grid and the cathode.

NIST ASD Team, NIST Atomic Spectra Database (version


3.1.5), [Online]. [Link] (2008)

used in the Franck-Hertz experiment.


[8] A. C. Melissinos, Experiments in Modern Physics Aca-
demic Press (1966).
[9] P.J. Mohr, B.N. Taylor, D.B. Newell, Rev. Mod. Phys.
FIG. 5: A blue-white plasma formed at the filament (bot- 80, 633 (2008)
tom of glass tube). This plasma was observed at a variety
of values for temperature, Vacc and Vf . The image shown
above was taken at a temperature of 140 C with Vacc =25V
and Vf =6.2V.
reach the lowest lying excitation energies for mercury.
An interesting e↵ect was observed for particular sets of
Vf , Vacc and temperature. While the accelerating volt-
age was ramped up, a white-blue plasma formed near
the filament in the tube. This plasma grew as the accel-
erating voltage was increased. When the plasma finally
attached itself to the filament, the current readings no
longer moved in a sinusoidal pattern and the glowing
discharge in the expected range for mercury excitation
stopped [Fig 5]. We call the value of Vacc at which the
plasma attached itself to the filament the onset voltage,

70 V.
Vonset .
58
Vonset was observed to be dependent on both Vf and
the temperature of the tube. While this phenomenon
was not studied exhaustively, there seemed to be a nega-
tive, linear dependence on the temperature and a positive
linear dependence on Vf (which, in turn, was roughly re-
lated to the electron density in the tube) [Fig 6]. Further
study is warranted to characterize and understand this
phenomenon.
In summary, experimental results of the historic
increased from 0.0 to 45.0V in increments of 0.05V with 6.3V. Altering the temperature did not result in any sta-
a scan rate of 10 scans per second. For each run, Iag tistically significant change to the calculated excitation.
and Vacc were stored in a text file. Each file was then However, data analysis showed that, for higher temper-
analyzed using Microsoft Excel 2003, looking at the de- atures, the current peaks began at a slightly lower volt-
pendence of Iag on Vacc [Fig 4]. Electrons that reached age. The maximum o↵set observed was 0.8V. This phe-
the anode were registered as negative current. Because nomenon was a result of the initial kinetic energy dis-
electrons were decelerated between the grid and anode, tribution of the electrons. For higher temperature, the
CHAPTER 3. ATOMS
this current is a measure of the electron energy. We have kinetic energy distribution is greater than that in a low59
reason to believe that the observed -2 nA o↵set visible in temperature environment. An electron with a large ini-
Fig 4 and all other data sets is a nonphysical product of tial kinetic energy required slightly less acceleration to
our data acquisition system.
The excitation potential for the favored 1 S0 !3 P1
Franck and Hertz noted that
transition was determined by multiplying the elementary
electron charge with the measured voltage di↵erence (for
the E = 4.9 eV characteris-
Vacc ) between the current peaks in a run. Data points
tic energy of the electrons in
representing the current peaks were selected by visually
inspection of the graphs produced by Microsoft Excel.
their experiment corresponded
For each run, the voltage di↵erence between adjacent
current peaks was calculated. The mean and standard
to the wavelengths λ = hc/E '
deviation of these measurements were calculated on a
run-by-run basis for all runs, including runs with varying
254 nm of light emitted by mer-
oven temperatures and filament voltages. The final ex-
1
citation potential was calculated by taking the weighted
cury atoms in gas discharges.
average of these measured values.
In fact, they also observed this
The excitation potential was measured to be 4.93 ±
0.06 eV. This value agrees with the expected value of
emitted light radiated from the
4.86 eV corresponding to the 1 S0 !3 P1 transition [7].
This determination of the excitation potential of mer-
tube in all directions of a single FIG.
cury also gives us an estimate for Planck’s constant, h.
4: Data from a run taken at 180 C with V at 6.3V. For
f
this run, the observed excitation potential was calculated to
We know
wavelength at 254 nm, but only the excitation be 4.89 ± 0.26eV. The arrows indicate the calculated value for
(3) Figure 3.8:
potential Anode
from all 6 datacurrent as a
sets, 4.93 ± 0.06eV,
hc/ = eV
if the voltage on the grid was big- typical run.
0 showing qualitative matching of the calculated value with this
where is the wavelength of the emitted light, c is the function of the voltage UGC be-
ger than 4.9 V.
tween the grid and the cathode as
They interpreted their dis-
measured in the Franck-Hertz ex-
covery as follows. Most collisions
periment.
between the mercury atom and
the electron are elastic. How-
ever, when the kinetic energy of the electron is about 4.9 eV the
collision with the mercury atom becomes inelastic, showing that the
electron could lose only a well defined (kinetic) energy before flying
away, leaving behind an excited mercury atom. A short time later,
the excited mercury atom released the deposited energy in the form
of ultraviolet light that has a wavelength of precisely 254 nm. Follow-
ing light emission, the mercury atom returns to its original, unexcited
state. Thus Franck-Hertz experiment, first conducted in 1914, was a
historic experiment that showed quantized internal energy excitation
in atoms.

Exercises
Atomic spectra

1. Compute the value of the ubiquitous parameter ρ0 = e2 /4πε0


where e is the unit charge.
2. An electron is moving on a circular orbit of radius r = 100 pm
around a (fixed) nucleus (for numerical estimates assume a pro-
ton). Find
(i) the total energy E (in units of eV)
1 They attributed this formula to J. Stark and to A. Sommerfeld, but we al-

ready know that the same formula was used by Einstein in his explanation of the
photoelectric effect in 1905.
60 CHAPTER 3. ATOMS

(ii) the frequency of the classical oscillation (in units of 1/fs),


(iii) the radiated power P as a function of the energy E and
(iv) the total radiated energy in one full cycle (E) in units of
the modulus of the energy |E|
of the electron.
(v) Give your answers to questions (i–iv) also when r = 1 pm.
3. An electron is orbiting a (fixed) proton of radius rn ' 1 fm at a
distance r0 = 100 pm (= 1 Angström). Compute the lifetime τ
of such a classical atom (the time needed for the electron to fall
into the nucleus).
4. Compute the energy of the lowest energy level (ground state),
called Rydberg energy, of the hydrogen atom.
5. In a thought experiment we illuminate a prism made of atomic
hydrogen with a beam of laser light of wavelength λ = 434.1 nm.
What do you expect to see?
6. Test the correspondence principle for states with n = 10i for
i = 0, 1, 2, 3 and 4 by computing the relative difference between
the frequencies of the classical and quantum theory.
7. In the lecture we assumed that the nucleus is fixed in space
because it is much more massive than the electron. Estimate
the correction to the Balmer-Rydberg formula if we take into the
possible motion of the proton in the hydrogen atom. The ratio of
the mass of the proton to that of the electron is mp /me ' 1860.

3.2.2 Quantized angular momentum


The quantized energy levels are not so surprising as those were built
in by the second postulate. As a result however, all other physical
characteristics of the electron will be quantized. For instance, from
Eq. (3.2.5), the angular speed is also quantized if the energy is quan-
tized: s
8En3
ωn = 2 .
ρ0 Z 2 m
More interesting quantity is the angular momentum
s r
2 8mr4 E 3 m
L = mr ω = − 2 2 = ρ0 Z −
ρ0 Z 2E
CHAPTER 3. ATOMS 61

of the electron. With quantized energy levels the momentum is also


quantized according to the following simple formula:
s
~2
r
m ma0
r
Ln = ρ0 Z − = ρ0 Z 2
n = ρ0 Z 2 n = ~n , (3.2.10)
2En ρ0 Z ρ0 Z 2

which predicts the angular momentum of the electron inside the atom
as a natural number times the natural unit ~. The question is whether
or not this prediction is supported by observations. The answer is a
clear ‘no’. For instance, in the ground state of the hydrogen atom
n = 1, so the predicted value for the angular momentum is L1 = ~,
while the measured value for the (orbital) angular momentum is 0.
Hence, we conclude that while Bohr’s postulates can predict the en-
ergy levels fairly precisely2 , the understanding of the quantized states
of the angular momentum requires a new theory. Also, the explana-
tion for the existence of stationary states requires clarification.

3.3 Splittings of spectral lines


No matter what this new theory is, it must also explain some other
observations that cannot be explained by Bohr’s postulates. We men-
tion three of those in connections with the atomic spectra: (i) the fine
structure of the spectral lines, (ii) the broadening of the spectral lines
in magnetic field and (iii) the broadening of the spectra in electric
field.

3.3.1 Fine structure


A.A. Michelson and E.W. Morley were the first to observe in 1887
that looking at the spectral lines of the hydrogen atom with sufficient
resolution, a single line is more often a pair of closely spaced lines.
This effect is small. For hydrogen like ions this splitting ∆En relative
to the Bohr energies En is proportional to (Zα)2 /n where α ' 1/137
is the fine structure constant. For instance, in the case of hydrogen
E21 ' 10.2 eV and the energy difference between the two components
of this transition from the first excited state to the two barely sepa-
rated levels in the n = 1 state is |∆E1 | ' 45 µeV, while in the case of
2 Later we shall see that even for the energy levels there are tiny differences

between the data and the predictions of Bohr.


62 CHAPTER 3. ATOMS

the famous sodium D-lines at 590 nm, the distance between the two
lines is about 0.6 nm, i.e. about 1 h. The derivation of the precise
formula  
En 2 1 3
∆En = (Zα) − (3.3.1)
n j + 12 4n

requires the understanding of the angular momentum j of the electron


in the atom.

3.3.2 Zeeman effect


In 1896 the Dutch physicist
P. Zeeman was the first who had
a sufficiently precise spectrome-
ter to observe the broadening of
spectral lines in magnetic field.
Furthermore, with higher fields
and better resolution he found
that a single line split some-
times into three lines (histori-
cally called normal Zeeman ef-
fect), while T. Preston of Ire-
land found in 1898 that it some-
times split into an even num-
ber of lines (historically called
anomalous Zeeman effect).
In the case of normal ef-
fect (shown in Fig. 3.9) Zeeman Figure 3.9: The spectral lines of
found that the spectral line char- mercury vapor lamp at wavelength
acterized by the frequency f0 in 546.1 nm. (A) Without magnetic
a magnetic field B, in the radia- field. (B) With magnetic field,
tion emitted perpendicularly to spectral lines split as transverse
the direction of the field, splits Zeeman effect. (C) With magnetic
into three lines with frequencies field, split as longitudinal Zeeman
f0 and f0 ± ωL /2π where ωL = effect.
−γB is called the Larmor fre-
quency. (The lines in the radiation emitted parallel to the B field
split into two.) The factor
gq
γ= (3.3.2)
2m
CHAPTER 3. ATOMS 63

is the gyromagnetic ratio for a particle of charge q and mass m, and


it is the ratio of the magnetic moment to the angular momentum of
the particle, so µ ~ The magnetic moment has SI unit [µ] = J/T,
~ = γ L.
but for particles the eV/T is more natural. The number g is called g
factor, first introduced by A. Landé in 1921.
For the classical loop current of a single electron moving on a circle,
we can compute easily that
e ~
~ =−
µ L, (3.3.3)
2me
e
so γe−loop = − 2m e
and the corresponding g factor is one. Note that
γ depends only on the charge and mass of the particle, showing that
the magnetic moment and angular momentum of a particle cannot be
separated. In fact, we can never observe the angular momentum of
particles, only their magnetic dipole moment. This was also demon-
strated by an experiment carried out by Einstein and W.J. de Haas
in 1915.
In the class we performed the demonstration of the effect. We suspended an
iron bar from a thin wire. The bar was surrounded with a solenoid. Switching on
and off periodically a current through the solenoid we observed a rotation of the
bar. How can we interpret this effect?
It was already known that iron is a ferromagnet, which were known
to be made up of a number of magnetic domains. These are regions of
the crystal throughout which the atomic magnetic dipoles are aligned,
but the domains are not aligned. Instead, those are oriented randomly,
such that their individual contributions to the total external magnetic
field largely cancel with one another. When the external homogeneous
field of the solenoid appears due to the electric current, then the mag-
netic field of the domains line up in the direction of the external field.
At the same time the bar turns, i.e. its angular momentum changes
from zero to some non-vanishing value. In fact, we know that the
direction of this angular momentum is the same as the direction of
the external magnetic field.
A natural explanation of the effect based on electron angular mo-
menta, and also the conclusion of Einstein and de Haas is the follow-
ing. In a domain the magnetic momenta of the electrons are aligned
adding up to the magnetic moment µ ~ d of the domain. The magnetic
potential energy of such a dipole in the external field B ~ is

U (θ) = −~ ~ = −µd B cos θ


µd · B
64 CHAPTER 3. ATOMS

where θ is the angle between the directions of the domain and field.
This potential energy has its minimum when θ = 0, so the field acts
with a torque on the magnetic dipole momenta such that it turns their
directions to line up with the direction of the field. Then it follows
that the angular momenta belonging to the magnetic momenta also
line up, but opposite to the direction of the external field. In or-
der that the total angular momentum of the bar remains unchanged
(in the absence of external torque acting on the bar), the bar must
turn such that it has macroscopic angular momentum in the direction
of the magnetic field. Thus the Einstein-de Haas exepriment demon-
strates in a macroscopic way that atoms carry magnetism and angular
momentum in a non-separable way.
As the unit of angular momentum is the same as the unit of
Planck’s constant, we can also write the magnetic moment of the
electron-loop current as µ ~
~ e = −µB L/~ where

e~ J µeV
µB = ' 9.3 · 10−24 ' 57.9
2me T T
is called Bohr magneton, the natural unit of the magnetic moment in
the microworld.
There are lines that split into more than three lines. In such cases
the number of Zeeman sub-levels is even and the phenomenon is called
anomalous.
Zeeman shared the Nobel prize with H.A. Lorentz “in recognition
of the extraordinary service they rendered by their researches into the
influence of magnetism upon radiation phenomena” in 1902. The cor-
rect interpretation of the splitting of spectral lines emitted by atoms
in magnetic field required a long time and became possible essentially
with the birth of quantum mechanics.

3.3.3 Stern-Gerlach experiment


In 1922 O. Stern and W. Gerlach performed an experiment with stun-
ning result. You might remember the name of Stern for measuring
the Maxwellian distribution of speeds of molecules where he was an
expert. In this experiment he also experimented with a beam of silver
atoms that travelled through an inhomogeneous magnetic field. If the
silver atoms had magnetic moments, then the magnetic field would
deflect them. The force acting on the moments is F~ = grad~ ~
µ · B.
Although the field is inhomogeneous, the net force has (almost) only
CHAPTER 3. ATOMS 65

vertical component, so
∂B
F~ ≈ ~kµz .
∂z
They expected that the magnetic moments of the silver atoms were
pointing in all directions in space, hence they would be deflected into
a continuous line, depending on the size of their µz component.
The schematic setup of the
experiment is shown in Fig. 3.10
together with the expected and
measured result. They observed
two spots, which was a puzzle
not only for them, but for the
whole physics community. One
might suspect that the silver
atoms are complex, so in order
Figure 3.10: SternGerlach experi-
to exclude that the effect had
ment: a beam (2) of silver atoms
anything to do with that com-
emerging from an oven (1) and
plexity, in 1927 T.E. Phipps and
travelling through an inhomoge-
J.B. Taylor reproduced the effect
neous magnetic field (3) are de-
using hydrogen atoms in their
flected depending on the direc-
ground state. Stern and Gerlach
tion of their magnetic moment.
could measure the separation be-
The classically expected result is
tween the spots (about 0.2 mm)
a blurred line (4), but instead two
and thus deduce the force acting
well separated spots are observed
on the magnetic moments of the
(5).
silver atoms. Knowing the mag-
nitude of the magnetic field (it
was 0.1 T), they could conclude about the size of the magnetic mo-
ments belonging to the atoms at the two spots,

µz ≈ ±µB ~ , (3.3.4)

suggesting corresponding angular momenta for the atoms ±~. When


the experiment was repeated with hydrogen, one could think that
this was the angular momentum of the single electron in its ground
state, n = 1. In principle, the nucleus could also contribute, but as
it is confined to the centre of the atom it was not considered likely
that it has comparable angular momentum. As Bohr’s postulates
suggested the same value (see Eq. (3.2.10) one could think that this
was a confirmation for Bohr’s theory.
66 CHAPTER 3. ATOMS

Bohr’s theory was debated within the community seriously. Sup-


porters like A. Sommerfeld tried to extend it to take into account the
fine structure of the spectral lines. (In fact, Eq. (3.3.1) was obtained
by Sommerfeld correctly.) He assumed that the electrons move on
ellipses around the nucleus when the position is characterized by two
coordinates, the distance r from the nucleus and an azimuth φ, lead-
ing to one more quantum condition and quantum number k (besides
n). Without going into the details of the theory, we mention that the
azimuth quantum number3 k can take n integer values, k = 1, 2, . . . ,
n, which determine the shapes of the ellipses that all belong to the
same value of the energy determined by n.
An ellipse can be oriented in space in any directions, so to char-
acterize the position of the electron we need yet another angle θ.
Sommerfeld introduced a new quantum condition with new quantum
number m also for this angle. He found that m could take 2k + 1
integer values between −k and k. The values of different m belong to
the same value of k, hence to the same value of n. We say that the
energy levels are degenerate and an external force, such as magnetic
field can lift the degeneracy, resulting in splitting of the spectral lines.
Thus Sommerfeld’s extensions of Bohr’s quantization condition pro-
vided a natural explanation for the normal Zeeman effect (odd number
of lines), but could not explain the anomalous effect with even number
of lines (which was the origin of the adjective ‘anomalous’). Also it
did not explain why the silver beam in the Stern-Gerlach experiment
splits into two, instead of an odd number of beams.
Puzzled by the result of the Stern-Gerlach experiment, one might
think of the electron as a small charged, rotating ball, as suggested
first by R. Kronig and independently by S. Goudsmit and G. Uhlen-
beck in 1925. Under the influence of Pauli, Kronig did not publish
this idea, but the latter two did, so today we attribute the introduc-
tion of spin to Goudsmit and Uhlenbeck. If the electron is a tiny
rotating ball, then it has its intrinsic angular momentum, called spin
and denoted by S, and the corresponding magnetic dipole moment of
the electron is µ ~ As the gyromagnetic ratio does not depend
~ e = γe S.
on the radius of the loop of a classical loop current (see Eq. (3.3.3)),
the same result is valid for a classical rotating charged ball because it
can be divided into infinitesimal loop currents, each having the same
form of contribution to the magnetic moment. Then simple integra-
tion tells that Eq. (3.3.3) is also valid for a classical rotating charged
3 In spectroscopy the term secondary quantum number l = k − 1 was used.
CHAPTER 3. ATOMS 67

ball. Goudsmit and Uhlenbeck could explain both the fine structure
and the anomalous Zeeman effect assuming a g factor of 2 for the
spinning electron.
If the g factor for the electron is 2, then its gyromagnetic ratio is
γe = e/me . Furthermore, we should write the magnetic moment found
in the Stern-Gerlach experiment instead of the formula in Eq. (3.3.4)
as
~
µz ≈ ±2µB . (3.3.5)
2
As a consequence the corresponding angular momentum for the elec-
tron in the ground state of the hydrogen atom in magnetic field
is ±~/2. History has proven the correctness of the assumption of
Goudsmit and Uhlenbeck, which lead to the correct quantum me-
chanical description of the angular momentum of the electron inside
the atom.

3.3.4 Stark effect


Inspired by the Zeeman effect, J. Stark studied the spectral lines of
atoms put in electric field in 1913. He was the first to observe the
splitting of the lines in such circumstances. The correct understanding
of the Stark effect also requires quantum mechanics.

Exercises
Atoms in magnetic field

1. Consider an electron moving in a circular orbit with angular


momentum L.~ Compute the magnetic dipole moment and the
gyromagnetic ratio of this classical loop current.
2. Consider a magnetic dipole moment with gyromagnetic ratio
γ placed into an oscillating homogeneous magnetic field B ~ =
B(cos(ωt), 0, 0). (i) Write the equation of motion and solve it to
find the motion of this moment. (ii) Consider the case of static
field in the limit ω → 0.
3. Compute the Larmor frequency of a classical loop current con-
sisting of a single electron placed in a homogeneous magnetic
field B = 0.1 T.
68 CHAPTER 3. ATOMS

3.4 X-ray spectrum


So far
40-6weXdealt
RAYS with
AN DtheTHatomic
E OR DE spectra
R I NG in
OFtheTHinfrared,
E E LE M E visible
NTS and 1237
ultraviolet ranges. If we irradiate heavy metals with energetic electron
beam in the keV energy range then the atoms emit electromagnetic
ng of the Elements
radiation in the form of X-rays, i.e. with wavelength in the several

tens of pm
solid copper or tungsten, range. This
is bombarded radiation
with electronsis characteristic to the elements and

Relative intensity
was firstrange,
n the kiloelectron-volt observed by D.G. Barkla
electromagnetic radia-in 1909 who was awarded the Nobel
prize in 1917 “for his discovery
d. Our concern here is what these rays can teach us of the characteristic Röntgen [X-ray]
b or emit them. Figure 40-13 shows the wavelength
radiation of the elements”. Continuous
duced when a beam of 35
Such keV electrons
a radiation can falls on a using an X-ray
be studied K β that has a
spectrumtube
a broad, continuous spectrum
schematic viewofdepicted
radiation in
on Fig.
which3.11, left. λElectrons emerge from a
min
of sharply defined wavelengths.
heated cathodeThe andcontinuous spec- by the electric field between the
are accelerated
different ways, which
anodeweand
nextcathode.
discuss
40-6 Xseparately.
RAYS AN D TH E OR DE R I NG 1237
30 OF40TH E50E LE60
M E NTS
70 80 90
Wavelength (pm)
um
nd the Ordering of the Elements Figure 40-13 The distribution by wavelength of
nuous x-ray spectrum of Fig. 40-13, ignoring for the Kα
d target, such as solid copper or tungsten, is bombarded with electrons the x rays produced when 35 keV electrons
ent peaks that rise from it. Consider an electron of
Relative intensity

tic energies are in the kiloelectron-volt range, electromagnetic radia- strike a molybdenum [Link] sharp peaks
collides (interacts)
x rays is emitted. with one
Our concern hereofisthe
whattarget
these atoms,
rays canasteach
in us and the continuous spectrum from which they
ytoms
losethat
an absorb
amount or of energy
emit !K, which
them. Figure will appear
40-13 shows as
the wavelength rise are produced by different mechanisms.
Continuous
on that
f the is radiated
x rays producedaway
when from
a beam theof site of the
35 keV collision.
electrons falls on a spectrum K β
m target.
erred toWe theseerecoiling
a broad, continuous
atom becausespectrumof ofthe
radiation on which
relatively λmin
posed two peaks of sharply defined [Link] continuous spec-
we neglect that transfer.)
e peaks arise in different ways, which we next discuss separately.
in Fig. 40-14, whose energy is now less than K0, may 30 40 50 60 70 80 90
Wavelength (pm)
h a X-Ray
ous target atom, generating a second photon, with a
Spectrum
Figure 40-13 The distribution by wavelength of
is electron-scattering
amine the continuous x-ray process
spectrum can continue
of Fig. until the
40-13, ignoring for the
the x rays produced when 35 keV electrons
tationary. All the photons generated by these colli-
the two prominent peaks that rise from it. Consider an electron of strike a molybdenum [Link] sharp peaks
c energy
uous K0 that
x-ray collides (interacts) with one of the targetFigure
spectrum. atoms, as3.11:
in and the continuous spectrum from which they
The electron may lose an amount of energy !K, which will appear as rise are produced by different
Electron-atom mechanisms.
scattering
f that spectrum in Fig. 40-13 is the sharply defined
of an x-ray photon that is radiated away from the site of the collision.
w which
energy the continuous
is transferred toThe spectrum
kineticatom
the recoiling does
energy not
Ek exist.
because the This
gained
of by the
relatively
ofsponds
the atom; tohere
a collision
neglect in
weelectron thatwhich it an
transfer.)
until incident
reaches theelectron
cathode can
ttered electron K ∆K
Ek0-–ΔE
nergy K0 in ainsingle
Fig. 40-14, whose energy
head-on collision is now
withlessathan K0, may
target
k
be controlled precisely by
nd collision with a target atom, generating a second photon, with a
the voltage Target
ergy appears as between the energy theof two
a single photon, whose
electrodes. Inuntil
thethe ex- atom
hoton energy. This electron-scattering process can continue
minimum possible x-ray wavelength — is found
approximately stationary. All the photons generated by these colli- from
periment we observe the X-rays emit-
part of the continuous x-ray spectrum. KEk0
hc from the anode. A typical X-ray
ted
minentKfeature
0 # hf #
of that spectrum
, in Fig. 40-13 is the sharply defined Incident
spectrum is shown in does
[Link]
3.11, (= ∆ K)
hf (=ΔE k)
l
length lmin, below which min the continuous spectrum [Link].
This electron
X-ray
wavelength corresponds It hasto aseveral
collision characteristics:
in which an incident (i)electron
there photon
K0 – ∆ K
initial kinetic energy K0 in a single head-on collision with a target
ntially all this hc
isappears
a continuous spectrum, above which Target
" min # energy (cutoff
(ii) there
as the energy of a single photon, whose
wavelength).
are several (40-23)
discrete linesfrom
and
atom
wavelength —K the
0
minimum possible x-ray wavelength — is found
(iii) the continuous
hc spectrum starts at K0
, wavelength
Figure 40-14 An electron of kinetic energy K0
aKof
ally independent 0 #
well hfdefined
the #
target
l min material. If we were
λmin ,tocalled
passing near
Incident
hf (= ∆ K)
an atom in the target
electron may gener-
16
X-ray
target to a copper target,
cutoff for example,
wavelength thatall
is features of
independent of
ate an x-ray photon, the electron losing part
Figure 3.12:photon
-13 would change excepthc the cutoff wavelength. of its energy in the process. The continuous
" min # (cutoff wavelength). (40-23)
K0 x-ray spectrum arises in this way.
Figure 40-14 An electron of kinetic energy K0
wavelength is totally independent of the target material. If we were to
K-shell vacancy jumps from the shell with n ! 2 (called the L shell), the emitted
15 radiation is the Ka line of Fig. 40-13; if it jumps from the shell with n ! 3 (called
the M shell), it produces the Kb line, and so on. The hole left in either the L or M
shell will be filled by an electron from still farther out in the atom.
In studying x rays, it is more convenient to keep track of where a hole is

Energy (keV)
created deep in the atom’s “electron cloud” than to record the changes in the
CHAPTER
10 3. ATOMS quantum state of the electrons that jump to fill that hole. Figure 69 40-15 does
exactly that; it is an energy-level diagram for molybdenum, the element to
which Fig. 40-13 refers. The baseline (E ! 0) represents the neutral atom in its
ground state. The level marked K (at E ! 20 keV) represents the energy of the
the 5metal of the anode, depends onlyatom
molybdenum on with a hole in its K shell, the level marked L (at E ! 2.7 keV)
represents the atom with a hole in its L shell, and so on.
Ek . The relation
Kα between Ek The and λ min
transitions marked Ka and Kb in Fig. 40-15 are the ones that produce the two
is very simple: L (n = 2) x-ray peaks in Fig. 40-13. The Ka spectral line, for example, originates when an elec-
Kβ Lβ Lα
M (n = 3) tron from the L shell fills a hole in the K [Link] state this transition in terms of what
N (n = 4) the arrows in Fig. 40-15 show, a hole originally in the K shell moves to the L shell.
0
Figure 40-15 A simplified energy-level Ordering hc
diagram for a molybdenum atom, showing λ the=Elements,
min (3.4.1)
In 1913, BritishEphysicist
the transitions (of holes rather than elec- k H. G. J. Moseley generated characteristic x rays for as
many elements as he could find — he found 38 — by using them as targets for
trons) that give rise to some of the charac-
electron bombardment in an evacuated tube of his own design. By means of a
teristic x rays of that element. Each
trolley manipulated by strings, Moseley was able to move the individual targets
horizontal line represents the energy of
the atom with a hole (a missing electron) in
which indicates that at the into the path
cutoff of an electron beam.
wavelength the He measured
total kineticthe wavelengths
energy of the emitted
the shell indicated. x rays by the crystal diffraction method described in Module 36-7.
of the electron is transfered into
Moseleythethenemitted
sought (and found) regularitiesEin
radiation, = hf
k these minas, he moved from
spectra
element to element
with fmin = c/λmin . We interpret the inemergence
the periodic table.
ofInthe
particular, he noted that if, for a given
continuous
spectral line such as Ka, he plotted for each element the square root of the frequency
spectrum such that the electron passing
f against the near
position of an atom
the element is decelerated
in the periodic table, a straight line resulted.
Figure 40-16 shows a portion of his extensive data. Moseley’s conclusion was this:
by the electric field of the target atom, and the energy lost in this
deceleration is emitted in the We formhave here a proof that there is in the atom a fundamental quantity, which
increases ofregular
by X-ray stepsradiation
as we pass from(see Fig. 3.12).
one element to the next. This quantity
Hence this radiation is called bremsstrahlung can only be the charge on the (the
centralGerman
nucleus. word for
“breaking radiation”). As a result of Moseley’s work, the characteristic x-ray spectrum became the uni-
versally accepted signature of an element, permitting the solution of a number of
Moseley plot
The origin of the discrete
lines in the X-ray spectrum is 2.5
Pd Ag
completely different. The po- Mo Ru
2.0 Zr
No, the atomic
number and the
sition of these lines depend on Y Nb

atomic mass are notthe material of the anode, hence


√f (109 Hz1/2)

the same. They are 1.5 Zn


distinct concepts in this is called the characteris- Ni
Fe Cu
atomic physics and Cr Co
chemistry: tic X-ray spectrum. In 1913 1.0
Ti
V
Mn
K Ca
Atomic Number (Z)
H.G.J. Moseley studied the po- Cl
Si
Definition: The atomic
number is the number
sition of the first line–called Kα
Figure 40-16 A Moseley plot of the K line of 0.5
Al
a
of protons in the the line–in thex-ray
characteristic spectra
spectra of of
21 heavy met-
nucleus of an [Link]. The frequency is calculated from
Significance: It als. When plotting
the measured wavelength. the square 0 10 20 30 40 50
uniquely identifies an Element number in periodic table
element. For example, root of the corresponding fre- atomic number
carbon has an atomic
number of 6, meaning
quency as a function of the po-
17
every carbon atom sition of the element in the peri-
has 6 protons.
Representation: It isodic table (the atomic number), Figure 3.13:
often denoted by the
symbol he found a straight line as can be
Z.
Atomic Mass (A) seen in Fig. 3.13. Before Moseley’s observation the ordering of the el-
Definition: The atomic
mass (also known as
ements in the periodic table had been according to their atomic mass,
atomic weight or mass although in some cases the chemical properties suggested that such
number) is the total
number of protons an ordering had to be inverted. As a result of Moseley’s plot, the
and neutrons in an
atom's nucleus. characteristic X-ray spectrum became the accepted means of identify-
Significance: It gives
an approximation of
ing a new element. Moseley himself suggested that the quantity that
the atom's mass. Foincreased in the atoms by fundamental steps was the electric charge

of the nucleus.
70 CHAPTER 3. ATOMS

3.4.1 Interpretation of Moseley’s plot


While Bohr’s postulates work well in explaining the hydrogen spec-
trum, it fails for atoms with more than one electron. In fact the
spectral lines in the visible spectrum do not show such regularity as
the characteristic X-rays. According to Eq. (3.2.9) the frequency of
the radiation emitted by the electron that jumps from the second en-
ergy level (n = 2) to the first one (n = 1) in a hydrogen-like atom
is  
13.6 eV 2 1 10.2 eV 2
f21 = Z 1− = Z . (3.4.2)
h 4 h
For large values of the atomic number, Z > 10, the corresponding
energy of the radiation is in the keV range, just like in the charac-
teristic X-ray spectrum. In fact, taking the square root of Eq. (3.4.2)
suggests a relation similar to Moseley’s plot. Thus we suspect that
the characteristic X-ray is emitted when an electron jumps between
the lowest energy levels. If we assume that the lowest energy levels
mean also the smallest distance of the electron from the nucleus, than
according to Gauss’ law the electron on such lowest energy levels does
not feel the electric force of the other electrons and behaves like the
electron of a hydrogen-like atom.
A closer look at Moseley’s plot reveals that the line crosses√ the
axis of the frequency at a negative value, suggesting a relation f21 =
a(Z − b) where
√ the two constants can √ be read off the plot: a ' (1.94 −
0.50) · 109 p Hz/(40 − 11) ' 49.6 THz and b ' 1. The value for a
agrees with 10.2 eV/h, supporting our interpretation. The value for
b can be understood if we assume that the lowest energy level can hold
two electrons and during the transition of the electron that emits the
characteristic radiation, there is another electron on the lowest energy
level, resulting in an effective charge (Z − 1)e seen by the radiating
electron.

3.5 Light amplification by the stimulated


emission of radiation
Lasers are wonderful examples of how practical devices of major im-
portance can appear as a result of attempts to solve scientific problems
that do not seem to have any relevance to technology. The concept
of stimulated emission first occurred to Einstein in 1917 when he was
CHAPTER 3. ATOMS 71

thinking about the blackbody radiation problem that he was able to


solve by uniting energy quantization and the photon concept of light.
By 1917 the following picture of an atom emerged. It consists of
a tiny, positively charged nucleus surrounded by electrons, so that
the atom appears neutral from large distances. The motion of the
electrons inside of the atom was not known, but it seemed that there
were states with different energy levels, characterized by quantum
numbers. Those levels were apparently stable, so that the electrons
could stay on those without emitting radiation. In the interaction of
matter with radiation the matter can absorb or emit radiation (called
absorption and spontaneous emission). In both cases the frequency
postulate (Eq. (3.2.6)) is valid. In absorption the radiation excites
the electron from a lower energy state to a higher one. In sponta-
neous emission the excited state has normally a short lifetime τ of the
order of 10 ns, during which the electron returns into the lower energy
state and emits radiation. Light from an old fashioned light-bulb is
generated by spontaneous emission. The photons emerging from the
bulb are totally independent: incoherent. In some cases the excited
state is called metastable because its lifetime is in the order of 1 ms.
There is a third way of interaction between matter and radiation.
In this case the electron is already in the excited state when the atom
is irradiated with radiation such that the frequency of the radiation
fulfills the frequency postulate. As a result, two identical photons are
emitted: the incoming one and the one emitted by the electron, both
with not only the same frequency, but also same direction, phase and
polarization. The process is called stimulated emission. Each of these
two photons can induce further stimulated emissions on neighbouring
excited atoms, leading to a chain reaction of stimulated emissions
caused by a single photon. Hence light becomes “amplified”.
Of course, the level of amplification depends on the number of
atoms in the excited state. According to the Boltzmann distribution,
for a large number of atoms in thermal equilibrium at a temperature T
the ratio of the number of atoms in the excited state x to the number
of atoms in the ground state 0 is

Ex − E0
 
Nx
= exp − .
N0 kB T
The typical value for the energy difference Ex − E0 is order of 10 eV,
while at room temperature kB T is about 25 meV, so the ratio in the
exponent is large, about 400, and the ratio is very small. According to
of N0 by of atoms between the ground state E
excited state Ex accounted for by th
N x $ N 0e#(Ex #E0)/kT, (40-29)
itation. (b) An inverted population,
in which k is Boltzmann’s constant. This equation seems reasonable. The quantity by special methods. Such a populati
kT is the mean kinetic energy of an atom at temperature T. The higher the sion is essential for laser action.
temperature,72 the more atoms — on average — will have been “bumped up” by 3. ATOMS
CHAPTER
thermal agitation (that is, by atom – atom collisions) to the higher energy state Ex.
Also, because Ex ! E0, Eq. 40-29 requires that Nx " N0; that is, there will always
be fewer atoms in the excited
Einstein, state than in
the probability of the ground state.
absorption of aThis is what
photon ofweenergy Ex − E0
expect if the level populations N0 and Nx are determined only by the action of
on an atom in its ground state
thermal agitation. Figure 40-19a illustrates this situation.
is the same as that of a stimulated
If we now flood the atoms of Fig. 40-19a with photons of energy Ex # E0, pho- radiate such
emission on an atom in its excited state. Hence, if we
tons will disappear via absorption
an ensemble by ground-state
of atoms atoms and
with photons ofphotons
energywill x −
Ebe gen-
E0 , most of the
W Discharge tube W
erated largelyphotons
via stimulated
will emission of excited-state
be absorbed by [Link]
their showed
groundthat state and much
the probabilities per atom for these two processes are [Link], because there
less will stimulate emission, simply because there are much Mmore of M2
are more atoms in the ground state, the net effect will be the absorption of photons. 1

To producethelaser
former
light,atoms
we mustthanhave the
morelatter ones.
photons emitted than absorbed; + Vdc –
(leak

If we
that is, we must have want that
a situation halfstimulated
in which of the atoms
emissionare in the Thus,
dominates. excited
we state, so that
Figure 40-20 The elements of a heliu
need more atoms in the excitedphotons
the irradiating state thancause
in the absorptions
ground state, as in Fig.
and 40-19b. emissions
stimulated inAn applied potenti
neon gas laser.
However, because
equal such a population
numbers, inversion is not
the temperature consistent
should be with thermal sends electrons through a discharg
equilibrium, we must think up clever ways to set up and maintain one. containing a mixture of helium gas
neon gas. Electrons collide with he
Ex − E0
The Helium–Neon Gas Laser T = ' 2 · 105 K . atoms, which then collide with neo
kB ln 2 which emit light along the length o
Figure 40-20 shows a common type of laser developed in 1961 by Ali Javan and
tube. The light passes through tran
his coworkers. The glass discharge tube is filled with a 20 : 80 mixture of helium windows
To produce laser light, we need even more photons emitted thanWab- and reflects back and f
and neon gases, neon being the medium in which laser action occurs. through the tube from mirrors M1
sorbed, so more atoms in the excited state than in the
Figure 40-21 shows simplified energy-level diagrams for the two types of atoms. ground one,
to cause more neon atom emission
calledpassed
An electric current population inversion,
through the which
helium – neon is clearly
gas mixture impossible
serves — through thermally
of the lightbe-
leaks through mirror M
collisions between
causehelium
the atoms
atomsanddisintegrate
electrons of the
atcurrent—to
much lowerraisetemperatures.
many helium form the laser beam.
The first successful experi-
mental demonstration of laser Metastable
state
light was performed by A. Ja-
The current (electrons)
van using a He-Ne gas laser. The
excite the helium atoms E3 E2
20
simplified diagram of the(but
by collisions energy
not the He–Ne
collisions E1 Laser light
levels of the twomore atoms is neon
massive shown (632.8 nm)
side by side in atoms).
Fig. 3.14. If
15 Excitation
an electron beam passes through Rapid Then the helium a
via collisions decay
the mixture of the gases, the excite the neon at
to level E2 by coll
Energy (eV)

electrons collide with the he-


10
Those neon atom
lium atoms and excite the elec- long enough to b
tron from its ground state to a into stimulated em
Figure 40-21 Fivemetastable
essential en- state of mean life-
ergy levels for helium and
5
neon atoms in time
a helium1–µsneonand at energy E3 '
gas laser. Laser20.61
action eV above the ground state
occurs
between levelsenergy
E2 and E1E of . The collision rate be-
0
neon when more atoms are at Common
the E2 level thantween the
at the E1
atoms of the gas mix- 0
Helium Neon
E0 ground state
level. ture is of the order of 1 billion states states
collisions per second, so the ex-
cited helium atom may collide
with a neon atom in the ground Figure 3.14:
state and excite its electron onto
a state with energy E2 ' 20.66 eV above E0 . This state 2 decays onto
CHAPTER 3. ATOMS 73

a state 1 with energy E1 ' 18.7 eV above E40-7 0 with a mean


L ASE RS lifetime
1243
of 170 ns. This latter state 1 has a much shorter lifetime. It decays
onto the ground state in about 10 ns. Thus we can achieve popula-
mulated emission tion forinversion:
a single atom.
state Suppose
2 will be more populatedExthan state 1 because Ex
ge number of atoms in thermal equilibrium at
(i) the metastability of state 3 and frequent collisions of the atoms
E0 E0
tion is directed at the sample, a number N0 of
ensures a continuous supply of neon atoms (a) in state 2; (ii)
(b) state 1 of
tate with energy E0 and a number Nx are in a
the neon atoms decay almost instantly. The photons emitted in the
g Boltzmann showed that Nx is given in terms 40-19 (a) The equilibrium distribution
spontaneous decay of state 2 causeFigure stimulated emission on neighbour-
of atoms between the ground state E0 and
ing neon atoms of state 2, resulting in a continuous
excited state Ex accounted redforlaser light ag-
by thermal of
$ N 0e#(Ex #E0)/kTwavelength
, λ = 632.8 nm. (40-29)
itation. (b) An inverted population, obtained
t. This equation seems Today lasers are
reasonable. Thecommonplace,
quantity byfor instance
special methods. everybody has seen
Such a population a
inver-
laser pointer. The
f an atom at temperature T. The higher the laser light sion
emerging is essential
from for
the laser action.
pointer is highly (i)
on average — will monochromatic
have been “bumped (as demonstrated
up” by by refraction on a prism), (ii) co-
– atom collisions) to the(as
herent higher energy state
demonstrated byEax. double-slit experiment), (iii) directional
requires that Nx(laser
" N0;light
that is,
is there
the bestwill realization
always of the ideal light ray) and also (iv)
ate than in thecan ground state. This
be focused is what
sharply. Atwe the European Laser Infrastructure (ELI)
and Nx are determined
the world only recordby of
theabout
action10 of25 W/cm2 flux density was achieved.
ustrates this [Link] laser light radiated from
Fig. 40-19a with photons
the He-Ne of energy x # E0, pho-
gas Emixture does Laser
by ground-statenot atoms and photons
immediately has will be gen-
these useful W Discharge tube W beam
ion of excited-state atoms. Einstein
properties becauseshowed that
the photons
two processes are [Link], because
are emitted from the neon atoms there
, the net effect will be the absorption of photons. M1 M2
in all directions. As a result, (leaky)
ust have more photons emitted than absorbed; + V dc –
most of the photons are stopped
which stimulated emission dominates. Thus, we Figure 40-20 The elements of a helium –
by the walls of the laser tube,
tate than in the ground state, as in Fig. 40-19b. neon gas laser. An applied potential Vdc
tion inversion whose schematic with
is not consistent picture is seen
thermal sends electronsFigure through3.15:
a discharge tube
in Fig. 3.15.
er ways to set up and maintain one. A voltage V dc accel- containing a mixture of helium gas and
erates electrons through a tube neon gas. Electrons collide with helium
containing a mixture of helium and neon
atoms, gases.
which Electrons
then collide excite
with neon the
atoms,
helium atoms, which then which
collide and emit
excitelightthe
alongneon
the length
atoms. of theThe
pe of laser developed in 1961 by Ali Javan and
latter emit tube. The
Thelight passes through transparent
e tube is filled with a 20 : 80laser lightofalong
mixture helium the tube.
windows W
light
and
passes
reflects
through
back and
trans-
forth
parent windows
dium in which laser action occurs. W and is reflected back and forth many times through
through the tube from mirrors M1 and M2
the tube
nergy-level diagrams from
for the twoconcave
types of mirrors
atoms. Mto 1 and M2 , with
cause more
(almost) coinciding
neon atom emissions. Some
focal
the helium – neon gaspoints
mixtureinserves
the middle
— through of the of
tube. Mirror M
the light leaks through
1 is coated
mirror M with
2 to
di-
d electrons of theelectric film whose
current—to raise manythickness
helium is chosen
form the such
laserthat
[Link] is almost totally
reflective at the wavelength of the laser light. The second mirror is
coated such that it is slightly transparent, so a fraction of the laser
Metastable
light escapes continuously, producing the useful laser beam.
state
e current (electrons)
cite the helium atoms 20 E3 E2
collisions (but not the He–Ne
collisions E1 Laser light
ore massive neon (632.8 nm)
oms).
74 CHAPTER 3. ATOMS

Exercises
X-rays and lasers

1. An electron beam falls on a manganese target anode in an X-ray


tube.
(i) What is the minimum value of the accelerating voltage V
that will produce Kα and Kβ lines of energies Eα = 5.899 keV
and Eβ = 6.491 keV? (ii) What is the cutoff wavelength for this
accelerating voltage? (iii) Compute λKα and λKβ .
2. Compute the ratio of wavelengths of the Kα lines in the spectra
of aluminium and silver.
3. We bombard a stainless steel target with electrons, and measure
the characteristic spectrum. We find a strong Kα line at λ1 =
1.936 Åand also a faint spectrum with Kα line at λ2 = 2.290 Å.
We know that stainless steel contains mostly iron, but what can
be the other element?
4. A He-Ne laser has an output power of P = 3 mW. Compute
the current density of the photons in the laser beam of cross
sectional area of A = 1 mm2 .
5. An atom has two energy levels with a transition wavelength
of λ = 582 nm. Due to population inversion there are 50 %
more atoms in the higher energy state than in the lower one.
What would be the corresponding temperature if the inverted
population were caused thermally?

3.6 Atomic quantum numbers


The experimental observations discussed in this chapter suggest that
the motion of electrons inside the atom is very different from the
classical motion of particles. In the latter case we can characterize
the motion of the particle precisely of we know the time-dependence
of its position vector. Applying the same concept to the electron inside
the atom, we found that the atom could not exist as a stable bound
system of the nucleus and the electrons. Bohr stated his first postulate
to resolve this apparent contradiction with reality, which however does
not explain why the electron can exist only on stationary states. The
CHAPTER 3. ATOMS 75

reason was found by Heisenberg, and his discovery opened the door
to a completely new description of physical reality, which will be the
subject of the course on quantum mechanics. In this section we shall
suffice with describing the characterization of the electrons according
to quantum mechanics without discussing the theory behind it.
Bohr’s second postulate lead naturally to the concept of quantized
energy levels of electrons inside the atom. We also mentioned that
it suggested the quantized nature of other kinematic and dynamical
quantities, such as angular momentum, but the predicted values were
not correct. We now present the correct values. The energy E is a
scalar quantity and single quantum number n is sufficient to charac-
terize it. Angular momentum L ~ is a vector, so it appears natural that
it requires more than one quantum number.

3.6.1 Magnitude of the angular momentum


We already mentioned that ~ is the natural unit of angular momentum
in the microscopic world, so we expect that the magnitude of the
angular momentum L is proportional to ~. Quantum mechanics tells
us that on an energy level characterized by n, L can take n discrete
values:
p
L = l(l + 1)~ , with l = 0, 1, 2, . . . , n − 1 (3.6.1)

where l is called the orbital quantum number. For instance, on the


ground state of the electron (n = 1) l = 0, so L = 0. In a hydrogen-
like atom for larger values of n the same energy belongs to more than
one values of orbital angular momentum, meaning different states of
motion. We say in such cases that the energy level is degenerate. In
atoms with more electrons this degeneracy is not present due to the
effect of the electric field of the other electrons in the same atom.

3.6.2 Direction of the angular momentum


~ we have to choose an arbitrary
When we talk about the direction of L,
direction with respect to which we measure the direction of the angular
momentum. Let us label that direction as the z axis. Quantum
mechanics tells us that for a given length of the angular momentum
characterized by l its z component can take discrete values only,

Lz = ml ~ with ml = 0, , ±1, ±2, . . . , ±l (3.6.2)


76 CHAPTER 3. ATOMS

where ml is called the magnetic quantum number. For l given, there


are 2l + 1 different, but degenerate states with the same orbital quan-
tum number and the same energy.
Let us consider a state with n = 2. According to Eq. (3.6.1), the
orbital quantum number can be l = 0, or 1. √ In the l = 1 state, the
magnitude of the angular momentum is L = 2~ ' 1.41~, while its
z component can be Lz = 0, or ±~. So only three values of Lz are
allowed. Furthermore, max Lz < L, which is independent of the value
of l and ml . Thus L ~ cannot be aligned with the z axis. This may
sound nonsense, as z was chosen arbitrarily, so L ~ cannot be aligned
with any direction. To understand the meaning of this statement,
we should remember that angular momentum and magnetic moment
cannot be separated. Thus, if we want to observe the direction of
the angular momentum, we can do so only by observing the direction
of the magnetic moment, which is possible only in external magnetic
field. Hence, z (or any other specified axis) means the direction of
the external magnetic field if it exists, and in this case the magnetic
moment of the electron is also quantized according to

µ = −ml µB ,

which explains the name “magnetic quantum number”.


If no external field is applied, than Lz cannot be measured. All
we know is that for l given, the electron can have in 2l + 1 different
states, characterized by the different values of ml , but we cannot know
the actual directions of the angular momentum without actually ob-
serving it. Although this may sound strange, it is actually the same
phenomenon as we discussed in the context of the double-slit exper-
iment with electrons. In that case a single electron passed through
both slits until we did not observe explicitly which one it passed, but
then the interference disappeared. In the present case a single electron
in the atom is a mixture of states with different ml values, all having
the same energy. Once we put the atom into external magnetic field,
we measure the angle between its magnetic moment and the external
field precisely and the degeneracy of the energy level disappears, the
spectral lines split. We say that the external field lifted the degeneracy.
If we write Heisenberg’s uncertainty principle in its angular form
by substituting ∆y = r∆φ and r∆py = ∆Lz , then we have ∆Lz ∆φ ≥
~ where φ is the angle of rotation around the z axis. We see the
knowing Lz precisely, i.e. with ∆Lz = 0, ∆φ must be infinite, i.e. we
cannot say anything about the direction of the angular momentum
CHAPTER 3. ATOMS 77

around the z axis. Thus, measuring Lz , i.e. fixing the value of ml


precisely, means that if we measure Lz again, we obtain the same
value, but at the same time we cannot say anything about the value
of Lx or Ly . The measurement of say Lx after the measurement of
Lz can lead to 2l + 1 different values corresponding to the value of l.
It is only a convention to say that ml characterizes the value of Lz ,
we could equally say that it characterizes the value of Lx or Ly .

3.6.3 Intrinsic angular momentum


In the previous section we stated that for l given ml can take 2l + 1
values. However, this result seems to be in contradiction with the
Stern-Gerlach experiment where we found that a beam of the hydro-
gen atom in the ground state split into two instead of an odd number
of beams. From 2l + 1 = 2, we would conclude the l = 1/2, which
was not among the permitted values for l. However, you might recall
that the assumption of Goudsmit and Uhlenbeck about the g factor
of the electron implied that in the ground state of the hydrogen atom
the angular momentum of the electron could take ±~/2. Then we can
resolve the contradiction if we assume that the angular momentum J~
of a particle is a sum of two contributions, the orbital momentum L ~
and the intrinsic angular momentum, or as commonly called the spin
~ J~ = L
S, ~ + S.
~ The quantization of the spin follows the same rule as
that of the angular momentum in Eq. (3.6.1),
p
S= s(s + 1)~ , (3.6.3)

with s = 12 for the electron. With the discovery of new particles, it was
found that they always have spin either half integer (most commonly
1
2 ) or integer values (most commonly 1). We call the particles falling
into the first class fermions and the particles with integer spin bosons.
The total angular momentum is the sum of the orbital one and
the spin. In the ground state of the electron the orbital angular mo-
mentum is l = 0, so its total momentum is equal to its spin. Just like
the components of L, ~ the component of spin in an arbitrary direction,
conventionally called z is

1
Sz = ms ~ with ms = ± (3.6.4)
2
where ms is called the spin quantum number. The magnetic moment
78 CHAPTER 3. ATOMS

associated with spin is


e ~ ~,
~ =−
µ S = −2µB S
m
which is also quantized with ms ,
e
µz = − ms .
m

The total angular momentum J~ = L


~ +S~ is characterized by the
quantum number j, p
J = j(j + 1)~ . (3.6.5)
It is natural that j can take the values 0, ± 21 , ±1,. . . . The summation
of the quantum numbers l and s however might appear surprising,

j = l ± s.

3.7 Atoms
The state of the electron inside the atom is characterized uniquely
with four quantum numbers: n, l, ml and ms . Without external
electric or magnetic field, the energy is determined uniquely by n.
One energy level with principal quantum number n is called a shell.
In quantum mechanics shells are denoted according to their principal
quantum numbers. In spectroscopy and chemistry the capital letter
K, L, M etc. are used such that K corresponds to n = 1, L to n = 2
and so on.
A shell has n states with n different values of the orbital angu-
lar momentum quantum number l. Multi-electron atoms states with
different l have different energies. There are 2l + 1 states with differ-
ent values of the magnetic quantum number ml belonging to each l,
which are said to belong to the same subshell. Subshells are denoted
usually with letters. The correspondence between the letters and the
orbital quantum numbers is shown in Table 3.1. We denote a subshell
by writing its principle quantum number followed by a letter charac-
terizing the orbital quantum number. For instance, the subshell 3p
has n = 3 and l = 2.
Finally, the spin quantum number ms can take two values for any
electron with values n, l and ml fixed. Thus, the total number of
CHAPTER 3. ATOMS 79

Table 3.1:

l 0 1 2 3 4 5
subshell s p d f g h

degenerate states belonging to the shell of principal quantum number


n is
n−1
n(n − 1)
X  
2 (2l + 1) = 2 2 + n = 2n2 .
2
l=0

We are now almost ready to build atoms. In order to do so, we


state three principles:
1. The quantum number principle The quantum numbers n,
l, ml and ms of hydrogen-like ions describe precisely the states
of any electrons in any atom.
2. The minimum energy principle Inside an atom in its ground
state the electrons occupy the lowest energy levels, filling shells
in increasing order of n and subshells in the order of increasing
l as these orders mean increasing energy.
3. The Pauli exclusion principle Two electrons inside the atom
cannot have all the same quantum numbers. (Actually this prin-
ciple has a wider ranger of validity, stated as: two fermions
cannot occupy a single state, independently of the system of
fermions being considered, or in other words, two fermions can-
not have all the same quantum numbers.)
Employing these principles, we can build an atom by first defining
the number of protons inside the nucleus, i.e. the value of Z. Then
we fill the shells in increasing order of n with electrons until the total
number of electrons equals Z, making the atom neutral. The first
shell can hold 2 electrons, the second 8, the nth can hold 2n2 .
We denote the electron configuration of an atom by recording the
number of electrons in each subshell as a superscript after the name
of the subshell. For instance, 1s2 means that the lowest energy shell
contains 2 electrons, i.e. it is filled completely, while 2p4 means that
four states out of the six on the 2p subshell are filled. If a subshell is
filled completely, then its total angular momentum is zero. Thus, the
80 CHAPTER 3. ATOMS

angular momentum of an atom is determined by the angular momenta


of the electrons on its open subshell.

3.8 Periodic table of elements


We can now understand the organizing principles behind the periodic
table of elements. This table was invented by Mendeleev in 18?? by
observing similarities in the chemical properties of various elements.
In brief, the law
We already mentioned that after the discovery of Moseley’s plot, the states that the
basis of ordering in the table became Elements with similar chemical square root of the
frequency of the
properties appear in columns. emitted X-ray is
approximately
It is the easiest to understand the properties of the noble gases in proportional to
the atomic
the righmost column. These contain 2 (helium), 10 (neon), 18 (argon), number:
36 (krypton), 54 (xenon) and 86 (radon) electrons. A quick counting of .}
the states on the subshells reveals that the electrons inside the atoms
of these elements fill subshells completely, called closed subshells. The
detailed electron configurations are as follows:
He 1s2
Ne 1s2 2s2 2p6
Ar 1s2 2s2 2p6 3s2 3p6
Kr 1s2 2s2 2p6 3s2 3p6 4s2 3d10 4p6
Xe 1s2 2s2 2p6 3s2 3p6 4s2 3d10 4p6 5s2 4d10 5p6
Rn 1s2 2s2 2p6 3s2 3p6 4s2 3d10 4p6 5s2 4d10 5p6 6s2 4f 14 5d10 6p6
its ckear in note
(It is not obvious why a subshell with larger values of n, but lower
value of l has lower energy than another subshell with lower value of
n and higher value of l, such as 4s vs. 3d. Such inversions can only be
understood in quantum mechanics.) With closed subshells only, these
atoms do not contain any electron relatively far from the nucleus, so
that it could more easily intaract with electrons of other neighbouring
atoms, making these elements chemically almost inert.
The next easiest column to understand is the first one. Apart from
the hydrogen, these have the same filled subshells as the inert gases
plus one electron. For instance, sodium has the electron configura-
tion 1s2 2s2 2p6 3s1 . The average distance of the electron on the 3s
shell from the nucleus is much larger than the electrons on the lower
shells, making this electron more likely to interact with the electrons
of neighbouring atoms. We call it the valence electron of the sodium.
We already mentioned that closed subshells have zero angular mo-
mentum (and magnetic dipole moment), so the angular momentum of
CHAPTER 3. ATOMS 81

the sodium atom is due to the spin of its valence electron (l = 0 on the
s state). The alkali metals in the first column are chemically active
because their valence electrons combine readily with atoms that have
a “vacancy” (of electron) in their outermost shell.
The prime examples for such vacancies are the elements in the
column VIIA, called halogens. For instance, chlorine has 17 electrons
in a configuration 1s2 2s2 2p5 , which means that it has a vacany or
“hole” on its 3p subshell. If a chlorine and a sodium atom come close,
the valence electron of the letter will occupy this hole, making a strong
bound between the towo atoms. Indeed, NaCl is a stable compound.
The electron configurations of the elements can be deduced sim-
ilarly. The elements in column IIA (alkali earth metals), such as
potassium, have two valence electrons, while those in column VIA
(gases and metalloids), such as oxygen have two holes, which results
in stable compounds like CaO. Of course, elements with one valence
electron and two holes can also form compounds, the most common
example being H2 O.
It is fairly simple to derive the electron configurations of the ele-
ments in the first three horizonthal periods. In the 4th and 5th period
the 3d subshell gets filled before the 4p, leading to the appearence of
ten transtion metals in each period. In the 6th and 7th periods after
the 6s subshell, the 14 states of the 4f subshell start to fill, leading to
the inner transition metals: the lanthanide series in the 6th and the
actinide series in the 7th period.
It is remarkable that we can understand so much about the atoms
based on a few assumptions (quantization of energy according to
Bohr’s postulates, quantization angular momentum and the three
principles stated above). To learn about the motion of electrons inside
the atom we need quantum mechanics.

Exercises
Atoms and the periodic table of elements

1. Compute the magnitude of the angular momentum for orbital


quantum number l = 3 and the values of the possible Ly com-
ponents.
2. There are two electrons with n = 2 in an atom. How many
states are possible for these two (indistinguishable) electrons if
we (i) forget, or (ii) take into account Pauli’s exclusion principle.
3. Find the electron configuration of the silver atom and explain
82 CHAPTER 3. ATOMS

the result of the original Stern-Gerlach experiment.


4. Consider the lithium atom, Z = 3. (i) What are the quantum
numbers of its electrons? (ii) What are the quantum numbers of
its highest energy electron in the first excited state of the atom?
5. Using the periodic table of elements find the electron configura-
tion of (i) carbon, (ii) iron, (iii) uranium.
6. “Hund’s rules” refers to a set of rules used to determine the
detailed electron configuration (specification of the states of a
subshell that are filled) of the ground state of a multi-electron
atom. In chemistry the first rule is often referred to simply as
Hund’s Rule: For a given electron configuration, the electron
configuration with maximum multiplicity has the lowest energy.
Try to find the detailed configuration of silicon employing this
rule.
Chapter 4

Basics of Quantum
Mechanics

A thorough understanding of quantum mechanics requires new math-


ematics (functional analysis), which is beyond the scope of these lec-
tures. Nevertheless, we introduce the basic concepts that will help
thinking in a quantum mechanical way, which is quite new, and hence
strange. We start with the simplest possible example of the electron
spin. In the lectures on quantum mechanics we shell consider more
elaborate examples, leading to a gradual habituation to this rather
unsusual subject. The advantage of the electron spin is that it can be
described in a finite (two) dimensional linear space, making possible
simple and explicit computations.

4.1 Predictions in quantum mechanics


Let us recall the consequences of the double-slit experiment performed
with electrons. We found the following:

1. The notion of classical path cannot be employed in the micro-


world.

2. The probability that the electron passes through the two slits
is equal to the sum of probabilities that it passes through the
single slits separately, P12 = P1 + P2 only if we observe which
83
84 CHAPTER 4. BASICS OF QUANTUM MECHANICS

slit it went through. If we do not make such an observation, then


P12 6= P1 + P2 .

3. In an ideal experiment, in which all initial and final states are


well defined, the probability of an event is given by the mod-
ulus squared of a complex number ϕ, called probability ampli-
tude (similarly as the intensity of light is given by the modulus
squared of the light wave).

4. If an event can occur in several exclusive ways (for instance,


the electron reaches the detector either through slit 1 or slit
2), then the probability amplitude is the sum of the probability
amplitudes belonging to the event occurring in the exclusive
ways (ϕ12 = ϕ1 + ϕ2 in the example). This is the principle of
linear superposition.

In quantum mechanics in order to analyse an experiment three


distinct steps are needed. (i) In the first one we describe the initial
state by a probability amplitude. (ii) In the second one we describe
the evolution of this amplitude in time. The time dependence of this
amplitude does not describe the time dependence of the events. Ac-
cording to quantum mechanics the latter is impossible by Nature (not
by our incomplete understanding). (iii) In the final step we perform
a measurement on the object, and we can predict the probabilities of
the outcomes of such a measurement from the probability amplitude.
The predictive power of quantum mechanics lies in computing those
probabilities. We call the second step quantum dynamics, and we are
not dealing with that here.
In order to perform the first step, we have to find those observ-
able physical quantities among which measuring the value of one does
not influence the value of the other. We call such quantities simul-
taneously measurable observables. We mention that the position and
momentum of a particle are not such observables because measuring
one as precisely as possible make the other undetermined, as stated
by the uncertainty principle. The set of simultaneously observable
quantities depends on the object under consideration. Yet it is a ba-
sic assumption in quantum mechanics that for any object there is at
least one (in general more than one) complete and maximal set of
simultaneously measurable quantities. For instance, in the case of
a free electron, its momentum p~, its mass m, its electric charge −e,
~ one component of its spin (in an arbitrary direction) Sz
its spin S,
CHAPTER 4. BASICS OF QUANTUM MECHANICS 85

and its lepton number L form a complete set of observables. We also


know that the same electron in the Coulomb field of the nucleus can
be characterized by different observables, namely with the quantum
numbers n, l and ml instead of the momentum in the former set. We
call a complete measurement the sequence of such observations that
assign a number to each simultaneously measurable quantity. In such
a case we say that the object is in a pure quantum state, which we
can represent by a unique probability function. Of course, there is no
guarantee that our measurement was indeed complete. It is a more
general case that we do not know the pure state of the object. In such
cases we can employ a statistical approach that we shall discuss later.

4.2 Mathematical model of quantum me-


chanics
In searching for the correct mathematical model of quantum mechan-
ics the principle of linear superposition will guide us. The linear com-
bination of two probability amplitudes is also a probability amplitude,
hence the quantum state should be described by a vector in an ab-
stract vector space. To understand the operations we need to define
on the elements of this vector space we return to the Stern-Gerlach
experiment.

4.2.1 Physical states as vectors


With the discovery of the electron spin, we can consider the Stern-
Gerlach experimental setup as an apparatus that can separate elec-
trons with their spin pointing up (Sz = ~/2) and pointing down
(Sz = −~/2). We call such a separator a Stern-Gerlach apparatus.
Fig. 4.1 shows such multiple Stern-Gerlach apparatuses (the rect-
angles containing S-G) linked one after the other. A shaded box rep-
resents a wall blocking the beam running into it. An S-G apparatus
can be oriented such that it splits beams into two beams containing
spins pointing up or down with respect to an arbitrary direction. We
usually choose this direction to be the z axis and we denote the spins
pointing up with a vector |z, +i and that pointing down with a vector
|z, −i. We can also split the beam into two beams containing spins
pointing to the right (denoted by |x, +i) or to the left (denoted by
|x, −i).
86 CHAPTER 4. BASICS OF QUANTUM MECHANICS

Figure 4.1:

In the upper figure, the first S-G apparatus splits the beam ac-
cording to the spins in the vertical direction, and then the spins in
the |z, −i are blocked. If we let the remaining beam through another
S-G apparatus, separating also in the z direction, then we do not
observe any more separation. This shows that the states |z, +i and
|z, −i have zero overlap, so these are orthogonal states, which can be
expressed formally as
hz, −|z, +i = hz, +|z, −i = 0 .
The operation in this equation is an inner product defined on the
elements of the abstract vector space.
In the 2nd S-G apparatus the intensity of the |z, +i beam does not
change, so we have perfect overlap between states |z, +i and |z, +i,
hz, +|z, +i = 1 (and also hz, −|z, −i = 1) .
We conclude that the states |z, +i and |z, −i form an orthonormal
basis of a two dimensional complex vector space. It cannot be a real
space because there |z, +i and |z, −i are anti-parallel, not orthogonal,
which makes spin so strange to us. All other spin states (spins in any
other direction) can be expressed as linear combination of these basis
states.
The orthonormal basis can be written in a more compact form if
we introduce the following notation for the basis vectors:
|z, +i ≡ |1i , |z, −i ≡ |2i , hence hi|ji = δij .
CHAPTER 4. BASICS OF QUANTUM MECHANICS 87

As the vector space is complex, the linear combinations of basis states


contain complex coefficients,

|si = s1 |1i + s2 |2i , si ∈ C . (4.2.1)

In mathematics the complex vector space equipped with an inner


product is called a Hilbert space, so the quantum states are represented
by elements of a Hilbert space. In the case of spin this Hilbert space is
two dimensional, but in general it can have arbitrary number of dimen-
sions, even infinite dimensions. In quantum mechanics the elements
of this Hilbert space are called ket vectors, or kets, while the elements
of the dual space, hs| are called bra vectors, or bras. Hence the name
of the inner product is braket, which maps the product of the Hilbert
space and its dual onto the complex numbers, h | i : H ∗ × H → C.
The inner product on a Hilbert space has the following properties
for any elements |xi and |yi of H:

1. it is conjugate symmetric, hx|yi = hy|xi .
2. it is linear in its second argument, hx|ay1 + by2 i = ahy|x1 i +
bhy|x2 i for any complex numbers a and b.
(
hx|xi > 0 if x 6= 0 ,
3. it is positive definite for y = x,
hx|xi = 0 if x = 0 .
p
The length or norm of a ket is ||x|| = hx|xi. The linear superposition
of a state vector with itself does not form a new state, hence the norm
of a ket is not observable. For the sake of simplicity, we assume that
its norm is equal to unity. However, the inner product of two different
state vectors has physical meaning that we do not define explicitly
here.
entangled In the middle figure we employ an S-G apparatus separating in
with the spin
state along the x direction (after the first one that separated the beam in the z
the
x-direction as
direction). What we observe is that the beam of |z, +i states contain
well. an equal number of |x, +i and |x, −i states. So a spin state along the
x direction has overlap with a spin state along the z direction,

hx, +|z, +i =
6 0, and also hx, −|z, +i =
6 0.

Hence the vectors |x, ±i are linear combinations of the basis vectors
|z, ±i,
(x,±) (x,±)
|x, ±i = c+ |z, +i + c− |z, −i .
88 CHAPTER 4. BASICS OF QUANTUM MECHANICS

In the third figure we first repeat the previous experiment and


select the beam containing the |x, +i states. As these were previously
filtered by a S-G apparatus in the z to contain a beam of |z, +i states
only, we might expect that employing yet another S-G separation in
the z direction gives us a beam of |z, +i states only. But this is not
the case. The second S-G apparatus filtering in the x direction erases
the memory of the first one and so the third apparatus yields beams of
both |z, +i and |z, −i states with equal intensity. We may say that an
S-G apparatus must alter the states of the particles that pass through
it, and the new state is a linear combination of basis vectors in the x
direction.
(z,±) (z,±)
|z, ±i = c+ |x, +i + c− |x, −i .

4.2.2 Representations of state vectors


Just like we use coordinates of the position vectors in the geometri-
cal space when we make explicit computations, it is often convenient
to use explicit representations of the abstract vectors of the Hilbert
space. The dimensionality of the Hilbert space depends on the actual
physical system, so these representations also depend on the system.
For instance, if the system can only have two basic states, and all other
states can be described by linear combinations of these basic states,
then the Hilbert space is two dimensional and the basis has two ele-
ments. An example for such a system is the double-slit experiment of
the electron when one basis state is that the electron passes through
slit 1 and the other is when it passes through slit 2. The other, more
important example is the electron spin with the two basis vectors |1i
and |2i. Then a general spin state is given by the linear combination
of Eq. (4.2.1), and the complex coefficients s1 and s2 are the prob-
ability amplitudes that represent the state uniquely on the basis |ii.
The modulus squares |si |2 are the probabilities that measuring the
spin of state |si results in a value Sz = ~/2 (for i = 1, pointing into
the direction of the magnetic field used in the S-G apparatus), or in
Sz = −~/2 (for i = 2, pointing opposite to the magnetic field). Thus
the general spin state can be represented by two complex numbers
that we organize conveniently into a doublet
 
s1
|si ↔ .
s2
CHAPTER 4. BASICS OF QUANTUM MECHANICS 89

The symbol ↔ expresses that it is a one-to-one representation of the


abstract vector. In the following we shall simply use the equality sign
for such a representation.
Using the linearity of the vector space, we can also write the rep-
resentation in the form
   
1 0
|si = s1 + s2 ,
0 1

so the basis vectors can be represented by the doublets


   
1 0
|1i = and |2i = .
0 1

Then the inner product can be computed as a matrix multiplication


 
s1
hs|si = (s1 , s2 )∗ = |s1 |2 + |s2 |2 = 1 . (4.2.2)
s2

4.2.3 Physical quantities as operators


The result of a measurement on a quantum state is a quantum number
belonging to the physical quantity measured. Recall the first line of
Fig. 4.1. A repeated measurement of the same quantity results in the
same number. So the measurement is an operation on the quantum
state that returns a quantum number and a corresponding state. We
call these eigenvalue and eigenvector of the operator that acts on
the elements of the Hilbert space. For instance, the S-G apparatus
measures the z component of the spin. We denote the corresponding
operator by Ŝz and formalize the statement as
~
Ŝz |z, ±i = ± |z, ±i . (4.2.3)
2
As a result, on an arbitrary spin state |si the result of the spin mea-
surement is
~ ~
Ŝz |si = s1 |z, +i − s2 |z, −i ,
2 2
which is the manifestation that the operators acting on the Hilbert
space are linear operators. The modulus squares |si |2 of the prob-
ability amplitudes inform us about the probability of measuring the
values +~/2 or −~ on the state |si. Once the measurement is done,
the state is found in one of the basis states |z, ±i according to the
90 CHAPTER 4. BASICS OF QUANTUM MECHANICS

returned eigenvalue, so a repeated measurement of the same physical


quantity returns the same value, as discussed in the analysis of the
measurements with the S-G apparatus.
If the eigenvectors of the operator form a basis of the Hilbert space,
then we call the corresponding physical quantity measurable. If O is
a measurable quantity, then any state |xi can be expressed as a linear
combination of the basis kets {|ei i}ni=1 of the operator Ô,
n
X
|xi = xi |ei i .
i=1

Then the measurement of O on state |xi results


n
X
Ô|xi = xi oi |ei i , (4.2.4)
i=1

with oi being the eigenvalue belonging to the eigenket |ei i.


For any operator Ô : H → H we can define an adjoint operator
Ô† using the equation

hy|Ô† |xi = hx|Ô|yi .
The result of a measurement is a real number, so the eigenvalues of
an operator of a measurable quantity have to be real. Then we find
that ∗ ∗
hx|Ô|xi = x∗ = x = hx|Ô|xi = hx|Ô† |xi ,
so Ô = Ô† , which means that Ô has to be a self-adjoint operator. The
z component of the spin is measurable with an S-G apparatus, so the
corresponding operator is self-adjoint.
Let us now assume that the kets |ei i and |ej i are eigenvectors of a
self-adjoint operator Ô, belonging to different eigenvalues oi and oj :

Ô|ei i = oi |ei i , Ô|ej i = oj |ej i , oj 6= oj .

oi :a coefficient
Then we can write the following sequence of equalities:
or an ith
∗ ∗ ∗ ∗
element
associated w. oi hej |ei i = hej |Ô|ei i = hei |Ô† |ej i = hei |Ô|ej i = o∗j hei |ej i = oj hei |ej i
In the context of matrix elements, the bra state is used to "sandwich" an op. O
inn. prod gives a = oj hej |ei i , to compute the matrix element of the op-O
complex number between 2 basis states ket state ej, ei
representing the
hence hej |ei i = 0, i.e. the eigenvectors of a self-adjoint operator that
overlap or similarity
between the two states.
belong to different eigenvalues are orthonormal (there norm is 1 ac-
cording to our convention). Then the multiplication of Eq. (4.2.4)
CHAPTER 4. BASICS OF QUANTUM MECHANICS 91

with the bra hx| yields


X X
hx|Ô|xi = xi x∗j oi hej |ei i = oi |xi |2 ,
i,j i

which is the average value of theDmeasurement


E results oi (recall that
the |xi |2 are probabilities). Thus Ô = hx|Ô|xi is the average value
x
of the physical quantity O on state |xi, which obviously has a physical
meaning being the result of measurements.

4.2.4 Representation of operators


The action of a linear operator  on a basis ket belonging to the
operator Ô results in another ket that can be expanded on the basis
of Ô, X
Â|ei i = ak,i |ek i .
j

Multiplying this equation with hej | and using the orthogonality con-
important dition of the ket vectors gives a complex number,
consequences in
quantum mechanics.
For example: hej |Â|ei i = aj,i .
Real Eigenvalues:
Hermitian operators Those numbers can trivially put into a matrix form, so linear operators
have real eigenvalues.
The eigenvalues can be represented by matrices on a basis, which is called the O-
correspond to the A Hermitian
possible measurement representation of the operator Â. If the operator is self-adjoint, its matrix is a
outcomes of the
observable associated
matrix representation has to be Hermitian. complex square
matrix that is
with the operator. The z component of the spin can have two states, so Ŝz acts on a equal to its own
conjugate
Orthogonal two-dimensional Hilbert space. Such an operator can be represented transpose. In
Eigenvectors: The other words, if A
eigenvectors by a 2 × 2 complex matrix. A general 2 × 2 complex matrix has eight is a Hermitian
corresponding to matrix, then A is
real
distinct eigenvalues of a parameters (four elements, each having a real and an imaginary equal to the

orthogonal to each
part).
Hermitian operator are The condition of Hermiticity means that the elements in the complex
conjugate of its
diagonal have to be real, while those in the off diagonal are com-
other. This orthogonality transpose:
property is essential for
the orthogonal plex conjugates of each other. Hence the matrix that represents the A = A^†,
decomposition of
quantum states. operator Ŝz has four real parameters, that we can write in the form
Conservation of  
d + c a − ib
Probability: The
conservation of Ŝz ↔ = aσ1 + bσ2 + cσ3 + d1 (4.2.5)
probability in quantum
a + ib d − c
mechanics relies on the
Hermiticity of operators,
where we have introduced the notations for the Pauli matrices
as it ensures that the
total probability of all      
possible outcomes 0 1 0 −i 1 0
sums to one. σ1 = , σ2 = , σ3 =
1 0 i 0 0 −1
92 CHAPTER 4. BASICS OF QUANTUM MECHANICS

and 1 is the 2×2 unit matrix. Eq. (4.2.5) shows that these for matrices
form a basis in the vector space of 2 × 2 Hermitian matrices.
In order to find the correct representation of the operator Ŝz , we
need specifiy the real coefficients a, b, c and d in Eq. (4.2.5) such that
Eq. (4.2.3) is fulfilled. To do so first we work out the actions of the
Pauli matrices on the doublet representations of the |ii basis vectors
easily:
    
0 1 1 0
σ1 |1i = = ,
1 0 0 1
    
0 1 0 1
σ1 |2i = = ,
1 0 1 0
    
0 −i 1 0
σ2 |1i = = −i ,
i 0 0 1
    
0 −i 0 1
σ2 |2i = =i ,
i 0 1 0
    
1 0 1 1
σ3 |1i = = ,
0 −1 0 0
    
1 0 0 0
σ3 |2i = =− .
0 −1 1 1
We see that the matrix S3 = ~2 σ3 has exactly the required action
on the basis vectors (cf. with Eq. (4.2.3)), so it provides a correct
representation of the Ŝz operator, Ŝz ↔ S3 .
We have emphasized before that the direction of the z axis is ar-
bitrary, or more precisely, we define it by the direction of the external
magnetic field. Let us now first fix our coordinate system and choose
the direction of the external magnetic field in the direction described
by the polar and azimuthal angles (ϑ, ϕ). Then the spin operator is
a self-adjoint operator that depends on these two angles Ŝ = Ŝ(ϑ, ϕ).
In order not to carry the physical dimension of the spin operator, we
introduce its dimensionless, rescaled version by
~
Ŝ(ϑ, ϕ) ≡ σ̂(ϑ, ϕ) . (4.2.6)
2
We would like to find the representation of the σ̂(ϑ, ϕ) matirx. As it
is a Hermitian matrix, the decomposition given in Eq. (4.2.5) applies
also here and we have to find the real coefficients a − d. Before you go
on you might challenge your intuition and try to find out how those
numbers are related to the angles ϑ and ϕ.
CHAPTER 4. BASICS OF QUANTUM MECHANICS 93

The operator σ̂(ϑ, ϕ) acts on the elements of a two-dimensional


  space, the kets |si that can be represented by the doublets
Hilbert
u
. Let us denote the eigenvectors of the operator with |s, ±i, so
d  

σ̂(ϑ, ϕ)|s, ±i = ±|s, ±i, represented by |s, ±i = .

We introduce the notations σ̂x , σ̂y and σ̂z for the rescaled operators
for the spin in the x, y and z directions. Measuring these quantities on
the states |s, ±i many times, we obtain the following average values
(components of the unit vector in the direction of the spin):
vec_n

hσ̂x i± = ± sin ϑ cos ϕ , hσ̂y i± = ± sin ϑ sin ϕ , hσ̂z i± = ± cos ϑ .

We now choose the eigenvectors |ii (i = 1, 2) of the σ3 matrix as


representation of the basis in the Hilbert space. Then we can compute
the average value of σ̂z explicitly, equation computes the average value of the operator $\hat{\sigma}z
$ for the state $|s+\rangle$ using the eigenvectors of the $\sigma_3$
matrix as
 the basis
 in the Hilbert
 space.
∗ ∗ 1 0
hσ̂z i+ = hs+ |σ̂z |s+ i = (u+ , d+ ) u+ − d+ =
0 1
= |u+ |2 − |d+ |2 = cos ϑ .
Interference occurs when the amplitudes of different
quantum states interfere either constructively or Operators in quantum mechanics represent
Using also the normalization condition,
destructively. The phase factor introduces a relative observables, such as the spin of a particle.
phase between different components of a quantum probabilities of observing the spin in The operator $\hat{\sigma}_z$ represents
state. When these components interfere, the relative th state 2 2 the Pauli spin matrix in the z-direction. The
phase can determine whether the interference is hs+ |s+ i = |u+ | + |d+ | = 1 average value of an operator is obtained by
constructive (enhancing the probability of certain taking the expectation value of the operator
outcomes) or destructive (suppressing the probability in a given state.
of certain outcomes). we obtain two equations for u+ and d+ , which we can solve easily:
cos term is the projection onto z axis If cos Eigenstates: Eigenstates are special states
probability amplitude of a system that satisfy a specific equation
is 1, it means the spin vector is completely phase factor
r
aligned with the z-axis, indicating a high 1 + cos ϑ iψ ϑ iψ
involving an operator. In this case, the
u =e
probability of finding the spin in the up state. + = e cos , eigenvectors $|i\rangle$ of the $\sigma_3$
If cos is 0, it means the spin vector is 2 2 matrix serve as the eigenstates of the spin
operator.
orthogonal to the z-axis, indicating no r (-): neg. dir. along axis
preference for the up or down state. If cos 1 − cos ϑ
i(ψ+χ) ϑ i(ψ+χ)
d =e
is -1, it means the spin vector is completely
+ =e sin , : Operators can be represented as
anti-aligned with the z-axis, indicating a high
probability of finding the spin in the down
2 2 matrices in a specific basis.
$u_+$ and $d_+$, which are coefficients
state. associated with the eigenvectors $|i\rangle$ of The phase factor e^(i(+)) represents the overall phase or
or in doublet form, the $\sigma_3$ matrix. th represent the complex coefficient associated with the state. It
probability amplitudes or the weights of the introduces a phase difference relative to the other phase
The (1 + cos )/2 factor corresponding basis vectors in the state $|s_+\ factor, e^(i). The phase factor e^(i) represents the overall
ϑ
 
scales the cosine term to be
iψ cos 2 phase associated with the state. It introduces a complex
between 0 and 1 |s i = e
+ iχ ϑ .coefficient that can affect interference effects and
(normalization factr) e sin 2 measurement outcomes. The specific value of
determines the phase relationship between different
components of the state.
As the basis kets are orthogonal, hs− |s+ i = 0, we find the other basis
The phase factor e^(i(+)) includes an additional phase
vector of σ̂(ϑ, ϕ) as factor . This introduces an additional phase difference
relative to e^(i). The specific value of further influences
ϑ the interference effects and measurement outcomes.
 −iχ 
e sin
|s− i = eiψ 2 .
cos ϑ2
In quantum mechanics, the square root of the probability is often referred to as the probability amplitude. The probability amplitude represents the
magnitude of a quantum state's contribution to a particular outcome or measurement result.
94 CHAPTER 4. BASICS OF QUANTUM MECHANICS

The overall phase drops out from averages, hence it cannot be mea-
sured and can be chosen at wish. We set ψ = 0 for simplicity. We do
not yet know the meaning of χ.
The σ̂i (i = 1 or 2) operators in σ3 representation can be given by
the 2 × 2 Hermitian matrices
   
x11 x12 y11 y12
σ1 = , σ2 =
x∗12 x22 ∗
y12 y22

where x11 , y11 , x22 , y22 are real numbers, representing the average
values of measuring the x and y components of the spin on Ŝz eigen-
states pointing up and down. We know from the S-G apparatus anal-
yses that those mean values all vanish as we can measure both +~/2
and −~/2 values with equal probabilities. We also know from the
S-G measurements that the eigenvalues of the σ1,2 matrices, i.e. the
solutions of the characteristic equations λ2 − |x12 |2 and λ2 − |y12 |2 ,
are ±1, so |x12 | = |y12 | = 1. Let us assume that

eiξ eiη
   
0 0
σ1 = and σ 2 = .
e−iξ 0 e−iη 0

Then
ϑ ϑ  −i(χ+ξ) 
hσ1 i+ = cos sin e + ei(χ+ξ) = sin ϑ cos ϕ ,
2 2
hence χ = ϕ − ξ, and similarly
ϑ ϑ  −i(χ+η) 
hσ2 i+ = cos sin e + ei(χ+η) = sin ϑ sin ϕ ,
2 2
so cos(χ + η) = cos(ϕ − ξ + η) = sin ϕ = cos ϕ − π2 , and η − ξ = − π2 .


We can choose one of the phases freely, then the other two phases are
fixed. By convention we set ξ = 0. Then η = − π2 and χ = ϕ. The
complete solution that reflects the observations made with the S-G
apparatus with this convention reads

cos ϑ2 sin ϑ2
   −iϕ 
e
|s+ i = , |s− i =
eiϕ sin ϑ2 cos ϑ2

and
x12 = eiξ = 1 , y12 = eiη = −i .
Thus we see that the matrix representation of the operator σ̂x in the
σ3 representation is the Pauli matrix σ1 , while that of σ̂y is σ2 . We
CHAPTER 4. BASICS OF QUANTUM MECHANICS 95

can also compute the representation of the σ̂(ϑ, ϕ) operator in the σ3


representation. From the solutions for the eigenvectors,

ϑ ϑ
|s+ i = cos |1i + eiϕ sin |2i ,
2 2
ϑ −iϕ ϑ
|s− i = cos |2i − e sin |1i
2 2
we obtain
ϑ ϑ
|1i = cos |s+ i − eiϕ sin |s− i ,
2 2
ϑ ϑ
|2i = e−iϕ sin |s+ i + cos |s− i .
2 2
Then we can compute the matrix elements of the operator σ̂(ϑ, ϕ) as

ϑ ϑ
h1|σ̂(ϑ, ϕ)|1i = cos2 − sin2 = cos ϑ
2 2
ϑ ϑ −iϕ
h1|σ̂(ϑ, ϕ)|2i = 2 cos sin e = sin ϑe−iϕ
2 2
ϑ ϑ
h2|σ̂(ϑ, ϕ)|1i = 2 cos sin eiϕ = sin ϑeiϕ
2 2
2 ϑ ϑ
h2|σ̂(ϑ, ϕ)|2i = sin − cos2 = − cos ϑ ,
2 2
or in explicit matrix form

sin ϑe−iϕ
 
cos ϑ
σij (ϑ, ϕ) =
sin ϑeiϕ − cos ϑ
= σ1 sin ϑ cos ϕ + σ2 sin ϑ sin ϕ + σ3 cos ϑ .

Using the definitions of the Pauli matrices, you can derive easily
the following general decomposition of their products:

σi σj = δij 1 + i
X
εijk σk , (4.2.7)
k

and with the help of Eq. (4.2.7), we find the following commutation
and anti-commutation relations

εijk σk , {σi , σj } = 2iδij 1 .


X
[σi , σj ] = 2i
k
96 CHAPTER 4. BASICS OF QUANTUM MECHANICS

We can use Eq. (4.2.6) to reintroduce the physical dimension of spin


and find the commutation relations among the components of spin,
X
[Si , Sj ] = i~ εijl Sk , for instance, [Sx , Sy ] = i~Sz , (4.2.8)
k

which define the (Lie) algebra of the angular momentum components.


The matrices Si (i = 1, 2, 3) are called the fgenerators of the algebra.
The algebra of linear operators differs from that of real numbers only
due to the non-commutativity of operator products as opposed to the
commutativity of number products.

4.3 Heisenberg’s uncertainty principle


The basis vectors of simultaneously measurable physical quantities A
and B coincide, otherwise we could not assign a well-defined quantum
number to both quantities, formalized by the equations

Â|ai bj i = ai |ai bj i , B̂|ai bj i = bj |ai bj i .

Then it follows that

(ÂB̂ − B̂ Â)|ai bj i = 0 ,

so the operators of simultaneously measurable quantities commute. We


state without proof that this statement can be reversed: commuting
operators have common basis.
If two operators do not commute, then the corresponding physical
quantities cannot be simultaneously measured with arbitrary preci-
sion. In a state |xi the standard deviation of the measurement of the
physical quantity O is defined as
s
 D E 2 
∆O = Ô − Ô .
x x

If Ô is self-adjoint, then
 D E  D E   D E  2
∆O2 = hx| Ô† − Ô Ô − Ô |xi = Ô − Ô |xi .
x x x
CHAPTER 4. BASICS OF QUANTUM MECHANICS 97

According to Schwarz’ inequality, for two self-adjoint operators O1


and O2
 2  2  D E  2  D E  2
∆O1 ∆O2 = Ô1 − Ô1 |xi Ô2 − Ô2 |xi
x x
D D E  D E E 2
≥ Ô1 − Ô1 Ô2 − Ô2 ,
x x
(4.3.1)
 D E   D E 
with equality if and only if Ô1 − Ô1 |xi ∝ Ô2 − Ô2 |xi.
x x
Any linear operator Ô can be written as a linear combination of two
self-adjoint operators Ŝ1 and Ŝ2 , Ô = 12 Ŝ1 + 2i Ŝ2 , with Ŝ1 = Ô + Ô†
and Ŝ2 = iÔ† −D iÔ.E We
 applyD suchE a decomposition to the linear
operator Ô1 − Ô1 Ô2 − Ô2 . The operator corresponding
x x
to Ŝ1 will be
 D E  D E   D E  D E 
Ŝ1 = Ô1 − Ô1 Ô2 − Ô2 + Ô2† − Ô2 Ô1† − Ô1 ,
x x x x

while for Ŝ2 we have


h D E  D E   D E  D E i
Ŝ2 = i Ô2† − Ô2 Ô1† − Ô1 − Ô1 − Ô1 Ô2 − Ô2
x
  h ix x x

= i Ô2 Ô1 − Ô1 Ô2 ≡ i Ô2 , Ô1

where we used that Ôi are both self-adjoint operators. Then the right
hand side of the inequality (4.3.1) can be written as
 D E  D E  2
Ô1 − Ô1 Ô2 − Ô2
x x
1 D E2 1 D E 2 cross terms that
= Ŝ1 + Ŝ2 +
4 4 add to zero
1 Dh iE 2
≥ i Ô2 , Ô1 .
4 x

Substituting into the right hand side of Schwarz’ inequality, we obtain


 2  2 1 Dh iE 2
∆O1 ∆O2 ≥ i Ô2 , Ô1 ,
4 x

or after taking the square root,


   1 Dh iE
∆O1 ∆O2 ≥ i Ô2 , Ô1 , (4.3.2)
2 x
98 CHAPTER 4. BASICS OF QUANTUM MECHANICS

which is the generalized form of Heisenberg’s uncertainty relation. As


an example, we can choose Ô1 = Ŝx and Ô2 = Ŝy . Substituting
Eq. (4.2.8) into Eq. (4.3.2) with this choice, we obtain in either of the
states |ii (i = 1 or 2)
   ~ D E  2
~
∆Sx ∆Sy ≥ Ŝz = .
2 i 2

The spin is the simplest quantum mechanical system as it can


have only two states that form the basis of a two-dimensional Hilbert
space. All other quantum mechanical objects exist in higher dimen-
sional Hilbert spaces, the infinite dimension is the general case. Ob-
viously, those require a more formal mathematical foundation. We
discussed the spin explicitly because all mathematical features can be
derived explicitly from the analyses of measurements with S-G appa-
ratus.

You might also like