Algebraic Combinatorics Lecture Notes
Algebraic Combinatorics Lecture Notes
Jeremy L. Martin
University of Kansas
jlmartin@[Link]
Copyright ©2010–2023 by Jeremy L. Martin (last updated August 23, 2023). Licensed under a Creative
Commons Attribution-NonCommercial-ShareAlike 3.0 Unported License.
Contents
2 Poset Algebra 40
2.1 The incidence algebra of a poset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40
2.2 The Möbius function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
2.3 Möbius inversion and the characteristic polynomial . . . . . . . . . . . . . . . . . . . . . . . . 46
2.4 Möbius functions of lattices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
2.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
3 Matroids 63
3.1 Closure operators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
3.2 Matroids and geometric lattices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64
3.3 Graphic matroids . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
3.4 Matroid independence, basis and circuit systems . . . . . . . . . . . . . . . . . . . . . . . . . . 68
3.5 Representability and regularity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
3.6 Direct sum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
3.7 Duality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
3.8 Deletion and contraction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81
3.9 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
2
5.5 Supersolvable lattices and arrangements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
5.6 Beyond real hyperplane arrangements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
5.7 Faces and the big face lattice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
5.8 Oriented matroids . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
5.8.1 Oriented matroid covectors from hyperplane arrangements . . . . . . . . . . . . . . . 134
5.8.2 Oriented matroid circuits . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136
5.8.3 Oriented matroids from graphs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138
5.9 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138
3
9.12 What’s next . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 240
9.13 Alternants and the classical definition of Schur functions . . . . . . . . . . . . . . . . . . . . . 241
9.14 The Murnaghan–Nakayama Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 243
9.15 The Hook-Length Formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 247
9.16 The Littlewood–Richardson Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250
9.17 Knuth equivalence and jeu de taquin . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 252
9.18 Yet another version of RSK . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 254
9.19 Quasisymmetric functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 255
9.20 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 257
4
Foreword
These lecture notes began as my notes from Vic Reiner’s Algebraic Combinatorics course at the University
of Minnesota in Fall 2003. I currently use them for graduate courses at the University of Kansas. They will
always be a work in progress. Please use them and share them freely for any research purpose. I have added
and subtracted some material from Vic’s course to suit my tastes, but any mistakes are my own; if you find
one, please contact me at jlmartin@[Link] so I can fix it. Thanks to those who have suggested addi-
tions and pointed out errors, including but not limited to: Kevin Adams, Nitin Aggarwal, Trevor Arrigoni,
Dylan Beck, Jonah Berggren, Lucas Chaffee, Matthew Chen, Geoffrey Critzer, Mark Denker, Souvik Dey,
Joseph Doolittle, Ken Duna, Monalisa Dutta, Josh Fenton, Logan Godkin, Bennet Goeckner, Darij Grinberg
(especially!), Brent Holmes, Arturo Jaramillo, Alex Lazar, Kevin Marshall, Dania Morales, George Nasr (es-
pecially!), Nick Packauskas, Abraham Pascoe, Smita Praharaj, John Portin, Enrique Salcido, Billy Sanders,
Tony Se, and Amanda Wilkens. Marge Bayer contributed the material on Ehrhart theory in §7.6.
5
Chapter 1
1.1 Posets
Definition 1.1.1. A partially ordered set or poset is a set P equipped with a relation ≤ that is reflexive,
antisymmetric, and transitive. That is, for all x, y, z ∈ P :
1. x ≤ x (reflexivity).
2. If x ≤ y and y ≤ x, then x = y (antisymmetry).
3. If x ≤ y and y ≤ z, then x ≤ z (transitivity).
We say that x is covered by y, written x ⋖ y, if x < y and there exists no z such that x < z < y. Two
posets P, Q are isomorphic if there is a bijection ϕ : P → Q that is order-preserving; that is, x ≤ y in P
iff ϕ(x) ≤ ϕ(y) in Q. It is easy to check that isomorphism is an equivalence relation. A subposet of P is a
subset P ′ ⊆ P equipped with the order relation given by restriction from P .
We will usually assume that P is finite. Sometimes a weaker assumption suffices, such that P is chain-finite
(every chain is finite) or locally finite (every interval is finite). (We will say what “chains” and “intervals”
are soon.)
Remark 1.1.2. A digraph (short for “directed graph”) consists of a set V of vertices and a set E of edges,
which have the form − → for v, w ∈ V . A directed acyclic graph (or DAG) is a digraph with no directed
vw
cycles, i.e., edge sets of the form {−
v−→ −−→ −−−−→ −−→
1 v2 , v2 v3 , . . . , vn−1 vn , vn v1 }. If P is a poset, then the digraph with edges
−
→
{xy | x < y} is a DAG. Conversely, for any DAG, the relation “x ≤ y if there exists a directed path
x → · · · → y with zero or more edges” is a partial order. Thus posets and DAGs are essentially equivalent.
Example 1.1.3 (Boolean lattices). Let [n] = {1, 2, . . . , n} (a standard piece of notation in combinatorics) and
let 2[n] be the power set of [n]. We can partially order 2[n] by writing S ≤ T if S ⊆ T . A poset isomorphic to
2[n] is called a Boolean lattice of rank n. We may also use 2S or BoolS for the Boolean lattice of subsets of
any finite set S; clearly BoolS ∼= Bool|S| .
6
123 123
12 12 13 23 12 13 23
1 2 1 2 3 1 2 3
∅ ∅ ∅
Bool2 Bool3
The first two pictures are Hasse diagrams: graphs whose vertices are the elements of the poset and whose
edges represent the covering relations, which are enough to generate all the relations in the poset by tran-
sitivity. (As you can see on the right, including all the relations would make the diagram unnecessarily
complicated.) By convention, bigger elements in P are at the top of the picture.
The Boolean lattice 2S has a unique minimum element (namely ∅) and a unique maximum element (namely S).
Not every poset has to have such elements, but if a poset does, we will call them 0̂ and 1̂ respectively (or if
necessary 0̂P and 1̂P ).
Definition 1.1.4. A poset that has both a 0̂ and a 1̂ is called bounded.1 An element that covers 0̂ is called an
atom, and an element that is covered by 1̂ is called a coatom. For example, the atoms in 2S are the singleton
subsets of S, and the coatoms are the subsets of cardinality |S| − 1.
We can make a poset P bounded: define a new poset P̂ by adjoining new elements 0̂, 1̂ such that 0̂ < x < 1̂
for every x ∈ P . Meanwhile, sometimes we have a bounded poset and want to delete the bottom and top
elements.
Definition 1.1.5. Let x, y ∈ P with x ≤ y. The interval from x to y is the set
[x, y] = {z ∈ P : x ≤ z ≤ y}.
This formula makes sense if x ̸≤ y, when [x, y] = ∅, but typically we don’t want to think of the empty set as
a bona fide interval. Also, [x, y] is a singleton set if and only if x = y.
Definition 1.1.6. A subset C ⊆ P (or P itself) is called a chain if its elements are pairwise comparable. Thus
every chain is of the form C = {x0 , . . . , xn }, where x0 < · · · < xn . The number n is called the length of the
chain; notice that the length is one less than the cardinality of the chain. The chain C is called saturated if
x0 ⋖· · ·⋖xn ; equivalently, C is maximal among all chains with bottom element x0 and top element xn . (Note
that not all such chains necessarily have the same length — we will get back to that soon.) An antichain is
a subset of P (or, again, P itself) in which no two of its elements are comparable.2
For example, in the Boolean lattice Bool3 , the subset3 {∅, 3, 123} is a chain of length 2 (note that it is not
saturated), while {12, 3} and {12, 13, 23} are antichains. The subset {12, 13, 3} is neither a chain nor an
antichain: 13 is comparable to 3 but not to 12.
1 This term has nothing to do with the more typical metric-space definition of “bounded”.
2 To set theorists, “antichain” means something stronger: a set of elements such that no two have a common lower bound. On
the other hand, combinatorialists frequently want to talk about antichains in a bounded poset, where the more restrictive definition
would be trivial.
3 It is very common to drop the braces and commas when writing subsets of [n]: it is easier and cleaner to write {∅, 3, 123} rather
7
123 123 123 123
12 13 23 12 13 23 12 13 23 12 13 23
1 2 3 1 2 3 1 2 3 1 2 3
∅ ∅ ∅ ∅
chain antichain antichain neither
One of the many nice properties of the Boolean lattice Booln is that its elements fall into horizontal slices
(sorted by their cardinalities). Whenever S ⋖ T , it is the case that |T | = |S| + 1. A poset for which we can
do this is called a ranked poset. However, it would be tautological to define a ranked poset to be a poset
in which we can rank the elements! The actual definition of rankedness is a little more subtle, but makes
perfect sense after a little thought, particularly after looking at an example of how a poset might fail to be
ranked:
1̂
z
y
x
0̂
You can see what goes wrong — the chains 0̂ ⋖ x ⋖ z ⋖ 1̂ and 0̂ ⋖ y ⋖ 1̂ have the same bottom and top and
are both saturated, but have different lengths. So the “rank” of 1̂ is not well-defined; it could be either 2 or
3 more than the “rank” of 0̂. Saturated chains are thus a key element in defining what “ranked” means.
Definition 1.1.7. A poset P is ranked if for every x, y ∈ P , all saturated chains with bottom element x and
top element y have the same length. A poset is graded if it is ranked and bounded.
In practice, most ranked posets we will consider are graded, or at least have a bottom element. To define a
rank function r : P → Z, one can choose the rank of any single element arbitrarily, then assign the rest of
the ranks by ensuring that
x ⋖ y =⇒ r(y) = r(x) + 1. (1.1)
It is an exercise to prove that this definition results in no contradiction. It is standard to define r(0̂) = 0
so that all ranks are nonnegative; then r(x) is the length of any saturated chain from 0̂ to x. (Recall from
Definition 1.1.6 that “length” means the number of steps, not the number of elements — i.e., edges rather
than vertices in the Hasse diagram.)
In some sources “ranked” means “graded”; it can also be used for the stronger condition that all maximal
chains with the same top element have the same length. For example, the poset shown below, which
satisfies Definition 1.1.7, is not ranked in this stronger sense, since {w, x, y} and {z, y} are both maximal
elements of Cy : Anyway, I have never known this distinction to be important in practice.
Definition 1.1.8. Let P be a ranked poset with rank function r. The rank-generating function of P is the
8
Order ideal (generators) Order filter (generators) Interval (endpoints)
The expansion of this polynomial is palindromic, because the coefficients are a row of Pascal’s Triangle.
That is, Booln is rank-symmetric.
More generally, if P and Q are ranked, then P × Q is ranked, with rP ×Q (x, y) = rP (x) + rQ (y), and FP ×Q =
FP FQ .
Definition 1.1.9. A linear extension of a poset P is a total order ≺ on the set P that refines <P : that is, if
x <P y then x ≺ y. The set of all linear extensions is denoted L (P ) (and sometimes called the Jordan-Hölder
set of P ).
Colloquially, an order ideal is a subset of P “closed under going down”. Note that a subset of P is an order
ideal if and only if its complement is an order filter. The order ideal generated by Q ⊆ P is the smallest
order ideal containing it, namely ⟨Q⟩ = {x ∈ P : x ≤ q for some q ∈ Q}. Conversely, every order ideal has
a unique minimal set of generators, namely its maximal elements (which form an antichain).
Example 1.1.11. Let {F1 , . . . , Fk } be a nonempty family of subsets of [n]. The order ideal they generate is
∆ = ⟨F1 , . . . , Fk ⟩ = {G ⊆ [n] : G ⊆ Fi for some i} .
9
These order ideals are called abstract simplicial complexes, and are the standard combinatorial models for
topological spaces (at least well-behaved ones). If each Fi is regarded as a simplex (i.e., the convex hull of a
set of affinely independent points) then the order-ideal condition says that if ∆ contains a simplex, then it
contains all sub-simplices. For example, ∆ cannot contain a triangle without also containing its edges and
vertices. Simplicial complexes are the fundamental objects of topological combinatorics, and we will have
much more to say about them in Chapter 6. ◀
There are several ways to make new posets out of old ones. Here are some of the most basic.
Definition 1.1.12. Let P, Q be posets.
• The dual P ∗ of P is obtained by reversing all the order relations: x ≤P ∗ y iff x ≥P y. The Hasse
diagram of P ∗ is the same as that of P , turned upside down. A poset is self-dual if P ∼ = P ∗ ; the map
realizing the self-duality is called an anti-automorphism. For example, chains and antichains are
self-dual, as is Booln (via the anti-automorphism S 7→ [n] \ S). Any self-dual ranked poset is clearly
rank-symmetric.
• The disjoint union P + Q is the poset on P ∪· Q that inherits the relations from P and Q but no others,
so that elements of P are incomparable with elements of Q. The Hasse diagram of P + Q can be
obtained by drawing the Hasse diagrams of P and Q side by side.
• The Cartesian product P × Q has a poset structure as follows: (p, q) ≤ (p′ , q ′ ) if p ≤P p′ and q ≤Q q ′ .
This is a very natural and useful operation. For example, it is not hard to check that Boolk × Boolℓ ∼ =
Boolk+ℓ .
• Assume that P has a 1̂ and Q has a 0̂. Then the ordinal sum P ⊕ Q is defined by identifying 1̂P = 0̂Q
and setting p ≤ q for all p ∈ P and q ∈ Q. Note that this operation is not in general commutative
(although it is associative).
P Q P ×Q P ⊕Q
1.2 Lattices
Definition 1.2.1. A poset L is a lattice if every pair x, y ∈ L has (i) a unique largest common lower bound,
called their meet and written x ∧ y; (ii) a unique smallest common upper bound, called their join and
written x ∨ y. That is, for all z ∈ L,
z ≤ x and z ≤ y ⇒ z ≤ x ∧ y,
z ≥ x and z ≥ y ⇒ z ≥ x ∨ y,
10
Note that, e.g., x ∧ y = x if and only if x ≤ y. Meet and join are easily seen to be commutative and
associative, so for any finite M ⊆ L, the meet ∧M and join ∨M are well-defined elements of L. In particular,
every finite lattice is bounded, with 0̂ = ∧L and 1̂ = ∨L. (In an infinite lattice, the join or meet of an infinite
set of elements may not be well-defined.4 ) For convenience, we set ∧∅ = 1̂ and ∨∅ = 0̂.
It is easy to see that any poset isomorphic to a lattice is a lattice, and meet and join are equivariant under
isomorphism (i.e., if f is an isomorphism, then f (x∧y) = f (x)∧f (y) and f (x∨y) = f (x)∨f (y)). (Therefore,
in order to show that two lattices are isomorphic, it is necessary only to show that they are isomorphic as
posets.)
The canonical example of a lattice is the Boolean lattice 2[n] . Its meet and join are intersection and union,
respectively. (In fact, the symbols ∧ and ∨ were probably chosen to resemble ∩ and ∪.)
Example 1.2.2 (The partition lattice). An [unordered] set partition of S is a set of pairwise-disjoint, non-
empty sets (“blocks”) whose union is S. It is the same data as an equivalence relation on S, whose equiva-
lence classes are the blocks. It is important to keep in mind that neither the blocks, nor the elements of each
block, are ordered.
Let Πn be the poset of all set partitions of [n]. For example, two elements of Π5 are
π = {1, 3, 4}, {2, 5} (abbr.: 134|25)
σ = {1, 3}, {4}, {2, 5} (abbr.: 13|4|25)
We can impose a partial order on Πn as follows: σ ≤ π if every block of σ is contained in a block of π; for
short, σ refines π (as here). To put it another way, σ can be formed by further splitting up π, or equivalently
every block of σ is a subset of some block of π. The lattices Π3 and Π4 are shown in Figure 1.3.
1234
1|2|3 1|2|3|4
Observe that Πn is bounded, with 0̂ = 1|2| · · · |n and 1̂ = 12 · · · n. For each set partition σ, the partitions
that cover σ in Πn are those obtained from σ by merging two of its blocks into a single block. Therefore,
Πn is graded, with rank function r(π) = n − |π|. The coefficients of the rank-generating function of Πn are
by definition the Stirling numbers of the second kind. Recall that S(n, k) is the number of partitions of [n]
into k blocks, so
Xn
FΠn (q) = S(n, k)q n−k .
k=1
4A lattice in which every set has a well-defined meet and join is called a complete lattice, although the concept does not arise often
in a combinatorial setting.
11
Furthermore, Πn is a lattice: any two set partitions π, σ have a unique coarsest common refinement
π ∧ σ = {A ∩ B : A ∈ π, B ∈ σ, A ∩ B ̸= ∅}.
Meanwhile, π ∨ σ is defined as the transitive closure of the union of the equivalence relations corresponding
to π and σ.
Finally, for any finite set, we can define ΠX to be the poset of set partitions of X, ordered by reverse
refinement; evidently ΠX ∼ = Π|X| . ◀
Example 1.2.3 (The connectivity lattice of a graph). Let G = (V, E) be a graph. Recall that for X ⊆ V , the
induced subgraph G|X is the graph on vertex set X, with two edges adjacent in G|X if and only if they are
adjacent in G. The connectivity lattice of G is the subposet of ΠV defined by
K(G) = {π ∈ ΠV : G|X is connected for every block X ∈ π}.
For an example, see Figure 1.4. It is not hard to see that K(G) = ΠV if and only if G is the complete graph
KV , and K(G) is Boolean if and only if G is acyclic. Also, if H is a subgraph of G then K(H) is a subposet
of K(G). The proof that K(G) is in fact a lattice (justifying the terminology) is left as an exercise.
1234
1 3
123|4 124|3 1|234 13|24
1|2|3|4 K(G)
◀
Example 1.2.4 (Partitions, tableaux, and Young’s lattice). An (integer) partition is a sequence λ = (λ1 , . . . , λℓ )
of weakly decreasing positive integers: i.e., λ1 ≥ · · · ≥ λℓ > 0. If n = λ1 + · · · + λℓ , we write λ ⊢ n and/or
n = |λ|. For convenience, we often set λi = 0 for all i > ℓ.
Partitions are fundamental objects that will come up in many contexts. Let Y be the set of all partitions,
partially ordered by λ ≥ µ if λi ≥ µi for all i = 1, 2, . . . . Then Y is a ranked lattice, with rank function
r(λ) = |λ|. Join and meet are given by component-wise max and min — we will shortly see another
description of the lattice operations. ◀
This is an infinite poset, but the number of partitions at any given rank is finite. In particular Y is locally
finite, i.e., every interval is finite.5 Moreover, the rank-generating function
X XX
q |λ| = qn
λ n≥0 λ⊢n
is a well-defined formal power series, and it is given by the justly celebrated formula
∞
Y 1
.
1 − qk
k=1
5 In general, if X is any adjective, then “poset P is locally X” means “every interval in P is X”.
12
There is a nice pictorial way to look at Young’s lattice. Instead of thinking about partitions as sequence
of numbers, view them as their corresponding Ferrers diagrams (or Young diagrams): northwest-justified
piles of boxes whose ith row contains λi boxes. The northwest-justification convention is called “English
notation”, and I will use that throughout, but a significant minority of combinatorialists prefer “French
notation”, in which the vertical axis is reversed. For example, the partition (5, 5, 4, 2) is represented by the
Ferrers diagram
(English) or (French).
Now the order relation in Young’s lattice is as follows: λ ≥ µ if and only if the Ferrers diagram of λ contains
that of µ. The bottom part of the Hasse diagram of Y looks like this:
In terms of Ferrers diagrams, join and meet are simply union and intersection respectively.
Young’s lattice Y has a nontrivial automorphism λ 7→ λ̃ called conjugation. This is most easily described
in terms of Ferrers diagrams: reflect across the line x + y = 0 so as to swap rows and columns. It is easy to
check that if λ ≥ µ, then λ̃ ≥ µ̃.
A maximal chain from ∅ to λ in Young’s lattice can be represented by a standard tableau: a filling of λ with
the numbers 1, 2, . . . , |λ|, using each number once, with every row increasing to the right and every column
increasing downward. The kth element in the chain is the Ferrers diagram containing the numbers 1, . . . , k.
For example:
1 2 4
∅ ⋖ ⋖ ⋖ ⋖ ⋖ ←→ .
3 5
Example 1.2.5 (Subspace lattices). Let q be a prime power, let Fq be the field of order q, and let V = Fnq (a
vector space of dimension n over Fq ). The subspace lattice LV (q) = Ln (q) is the set of all vector subspaces
of V , ordered by inclusion. (We could replace Fq with an infinite field. The resulting poset is infinite,
although chain-finite.)
The meet and join operations on Ln (q) are given by W ∧ W ′ = W ∩ W ′ and W ∨ W ′ = W + W ′ . We could
construct analogous posets by ordering the (normal) subgroups of a group, or the prime ideals of a ring, or
the submodules of a module, by inclusion. (However, these posets are not necessarily ranked, while Ln (q)
is ranked, by dimension.)
13
The simplest example is when q = 2 and n = 2, so that V = {(0, 0), (0, 1), (1, 0), (1, 1)}. Of course V has one
subspace of dimension 2 (itself) and one of dimension 0 (the zero space). Meanwhile, it has three subspaces
of dimension 1; each consists of the zero vector and one nonzero vector. Therefore, L2 (2) ∼ = M5 .
M5
Note that Ln (q) is self-dual, under the anti-automorphism W → W ⊥ (the orthogonal complement with
respect to any non-degenerate bilinear form).
The number of elements at rank k in Ln (q), i.e., the number of k-dimensional subspaces of Fnq , is the q-
binomial coefficient
(q n − 1)(q n − q) · · · (q n − q k−1 )
n
= ,
k q (q k − 1)(q k − q) · · · (q k − q k−1 )
The proof is left as an exercise (Problem 1.14(b)). For more on q-binomial coefficients (including a proof that
they are actually polynomials in q, not merely rational functions), see Problem 2.7. ◀
Example 1.2.6 (The lattice of ordered set partitions). An ordered set partition (OSP) of S is an ordered list
of pairwise-disjoint, non-empty sets (“blocks”) whose union is S. Note the difference from unordered set
partitions (Example 1.2.2). We use the same notation for OSPs as for their unordered cousins, but now, for
example, 14|235 and 235|14 represent different OSPs. The set On of OSPs of [n] is a poset under refinement: σ
refines π if π can be obtained from σ by removing zero or more separator bars. For example, 16|247|389|5 ≤
16|2|4|7|38|9|5, but 1|23|45 and 12|345 are incomparable. The Hasse diagram for O3 is as follows.
123
This poset is ranked, with rank function r(π) = |π| − 1 (i.e., the number of bars, or one less than the number
of blocks, just like Πn ). Technically On is not a lattice but only a meet-semilattice, since join is not always
well-defined. However, we can make it into a true lattice by appending an artificial 1̂ at rank n.
Interestingly, On is locally Boolean, i.e., every interval [π, σ] ⊆ On is a Boolean lattice, whose atoms corre-
spond to the bars that appear in σ but not in π.
There is a nice geometric way to picture On . Every point x = (x1 , . . . , xn ) ∈ Rn gives rise to an OSP
ϕ(x) that describes which coordinates are less than, equal to, or greater than others. For example, if x =
(6, 6, 0, 4, 7) ∈ R5 , then ϕ(x) = 3|4|12|5, since x3 < x4 < x1 = x2 < x5 . Let Cπ = ϕ−1 (x) ⊂ Rn ; that is,
Cπ is the set of points whose relative order of coordinates is given by π. Each set Cπ is a cone (i.e., it is
closed under addition and multiplication by positive scalars) and evidently the Cπ decompose Rn , so they
give a good picture of On . For example, the picture for n = 3 looks like this. (The picture is actually the
cross-section in the plane x1 + x2 + x3 = 0, but this cross-section is enough to see the full combinatorial
structure.)
14
x=z
x<z
x>z 1|3|2
123
3|1|2 1|2|3
13|2 1|23
x<y
3|12 12|3 x=y
x>y
23|1 2|31
3|2|1 2|1|3
y>z 2|3|1
y<z
y=z
The topology matches the combinatorics: for example, each Cπ is a |π|-dimensional space, and π ≤ σ in On
if and only if Cπ ⊆ Cσ (where the bar means closure). We will come back to this in more detail when we
study hyperplane arrangements in Chapter 5; see especially Example5.7.1. ◀
Example 1.2.7. Lattices don’t have to be ranked. For example, the poset N5 shown below is a perfectly
good lattice.
N5
◀
Proposition 1.2.8 (Absorption laws). Let L be a lattice and x, y ∈ L. Then x ∨ (x ∧ y) = x and x ∧ (x ∨ y) = x.
(Proof left to the reader.)
The following result is a very common way of proving that a poset is a lattice.
Proposition 1.2.9. Let P be a bounded poset that is a meet-semilattice (i.e., every nonempty B ⊆ P has a well-defined
meet ∧B). Then every nonempty subset of P has a well-defined join, and consequently P is a lattice. Similarly, every
bounded join-semilattice is a lattice.
Proof. Let P be a bounded meet-semilattice. Let A ⊆ P , and let B = {b ∈ P : b ≥ a for all a ∈ A}. Note
that B ̸= ∅ because 1̂ ∈ B. Then ∧B is the unique least upper bound for A, for the following reasons. First,
15
∧B ≥ a for all a ∈ A by definition of B and of meet. Second, if x ≥ a for all a ∈ A, then x ∈ B and so
x ≥ ∧B. So every bounded meet-semilattice is a lattice, and the dual argument shows that every bounded
join-semilattice is a lattice,
This statement can be weakened slightly: any poset that has a unique top element and a well-defined meet
operation is a lattice (the bottom element comes free as the meet of the entire set), as is any poset with a
unique bottom element and a well-defined join.
Definition 1.2.10. Let L be a lattice. A sublattice of L is a subposet L′ ⊆ L that (a) is a lattice and (b) inherits
its meet and join operations from L. That is,
Equivalently, a sublattice of L is a subset that is closed under meet and join. We can speak of the sublattice
of L generated by any subset S ⊆ L; it is just the smallest sublattice containing S.
Note that the maximum and minimum elements of a sublattice of L need not be the same as those of L. As
an important example, every interval L′ = [x, z] ⊆ L (i.e., L′ = {y ∈ L : x ≤ y ≤ z}) is a sublattice with
minimum element x and maximum element z. (We might write 0̂L′ = x and 1̂L′ = z.)
Example 1.2.11. Young’s lattice Y is an infinite lattice. Meets of arbitrary sets are well-defined, as are finite
joins. There is a 0̂ element (the empty Ferrers diagram), but no 1̂. On the other hand, Y is locally finite —
every interval [λ, µ] ⊆ Y is finite. Similarly, the set of natural numbers, partially ordered by divisibility, is
an infinite, locally finite lattice with a 0̂. ◀
Example 1.2.12. Consider the set M = {A ⊆ [4] : A has even size}. This is a lattice, but it is not a sublattice
of Bool4 , because for example 12 ∧M 13 = ∅ while 12 ∧Bool4 13 = 1. ◀
Example 1.2.13. [Weak Bruhat order] Let Sn be the set of permutations of [n] (i.e., the symmetric group).6
Write elements w ∈ Sn as strings w1 w2 · · · wn of distinct digits, e.g., 47182635 ∈ S8 . (This is called one-line
notation.) The weak Bruhat order ≤W on Sn is defined as follows: w ⋖W v if v can be obtained by swapping
wi with wi+1 , where wi < wi+1 . For example,
In other words, v = wsi , where si is the transposition that swaps i with i + 1. The weak order actually is a
lattice, though this is not so easy to prove.
Weak order is ranked by inversion number, and in fact v ≤W w if and only if I(v) ⊆ I(w).
Example 1.2.14. [Bruhat order] The Bruhat order ≤B on permutations is a related partial order with more
relations (i.e., “stronger”) than the weak order. It is defined as follows: w ⋖B v if inv(v) > inv(w) and v = wt
for some transposition t. For example,
47162835 ⋖B 47182635
6 That’s a Fraktur S, obtainable in LaTeX as \mathfrak{S}. The letter S has many other standard uses in combinatorics: Stirling
numbers, symmetric functions, etc. The symmetric group is important enough to merit an ornate symbol!
16
in Bruhat order (because this transposition has introduced exactly one more inversion), but not in weak
order (since the positions transposed, namely 4 and 6, are not adjacent). On the other hand, 47162835 is not
covered by 47862135 because this transposition increases the inversion number by 5, not by 1. (So this is a
relation, but not a cover, in Bruhat order.) ◀
The Bruhat and weak orders on S3 are shown below. You should be able to see from the picture that Bruhat
order is not a lattice.
321 321
123 123
A Coxeter group is a finite group generated by elements s1 , . . . , sn , called simple reflections, satisfying s2i = 1
and (si sj )mij = 1 for all i ̸= j and some integers mij ≥ 2. For example, setting mij = 3 if |i − j| = 1 and
mij = 2 if |i − j| > 1, we obtain the symmetric group Sn+1 . Coxeter groups are fantastically important in
geometric combinatorics and we could spend at least a semester on them. The standard resources are the
books by Brenti and Björner [BB05], which has a more combinatorial approach, and Humphreys [Hum90],
which has a more geometric flavor. For now, it’s enough to mention that every Coxeter group has associated
Bruhat and weak orders, whose definitions generalize those for the symmetric group.
The Bruhat and weak order give graded, self-dual poset structures on Sn , both ranked by number of in-
versions:
r(w) = {i, j} : i < j and wi > wj .
(For a general Coxeter group, the rank of an element w is the minimum number r such that w is the product
of r simple reflections.) The rank-generating function of Sn is a very nice polynomial called the q-factorial
(or “the q-analogue of n factorial”, “n factorial base q”, etc.):
n
Y 1 − qi
FSn (q) = 1(1 + q)(1 + q + q 2 ) · · · (1 + q + · · · + q n−1 ) = . ◀
i=1
1−q
Definition 1.3.1. A lattice L is distributive if the following two equivalent conditions hold:
x ∧ (y ∨ z) = (x ∧ y) ∨ (x ∧ z) ∀x, y, z ∈ L, (1.3a)
x ∨ (y ∧ z) = (x ∨ y) ∧ (x ∨ z) ∀x, y, z ∈ L. (1.3b)
Proving that the two conditions (1.3a) and (1.3b) are equivalent is not too hard, but is not trivial (Prob-
lem 1.8). Note that replacing the equalities with ≥ and ≤ respectively gives statements that are true for all
lattices.
17
The condition of distributivity seems natural, but in fact distributive lattices are quite special.
1. The Boolean lattice 2[n] is a distributive lattice, because the set-theoretic operations of union and in-
tersection are distributive over each other.
2. Every sublattice of a distributive lattice is distributive. In particular, Young’s lattice Y is distributive
because it is a sublattice of a Boolean lattice (recall that meet and join in Y are given by intersection
and union on Ferrers diagrams).
3. The lattices M5 and N5 are not distributive:
z
y a b c
x
(x ∨ y) ∧ z = 1̂ ∧ z = z (a ∨ b) ∧ c = c
(x ∧ z) ∨ (y ∧ z) = x ∨ 0̂ = x (a ∧ c) ∨ (b ∧ c) = 0̂.
4. The partition lattice Πn is not distributive for n ≥ 3, because Π3 ∼ = M5 , and for n ≥ 4 every Πn
contains a sublattice isomorphic to Π3 (see Problem 1.1). Likewise, if n ≥ 2 then the subspace lattice
Ln (q) contains a copy of M5 (take any plane together with three distinct lines in it), hence is not
distributive.
5. The set Dn of all positive integer divisors of a fixed integer n, ordered by divisibility, is a distributive
lattice (Problem 1.3).
Every poset P gives rise to a distributive lattice in the following way. The set J(P ) of order ideals of P (see
Definition 1.1.10) is itself a bounded poset, ordered by containment. In fact J(P ) is a distributive lattice:
the union or intersection of order ideals is an order ideal (this is easy to check) which means that J(P ) is a
sublattice of the distributive lattice BoolP . (See Figure 1.5 for an example.)
abcd
abc acd
b d ac cd
a c a c
P J(P )
For example, if P is an antichain, then every subset is an order ideal, so J(P ) = BoolP , while if P is a chain
with n elements, then J(P ) is a chain with n + 1 elements. As an infinite example, if P = N2 with the
product ordering (i.e., (x, y) ≤ (x′ , y ′ ) if x ≤ x′ and y ≤ y ′ ), then J(P ) is Young’s lattice Y .
Remark 1.3.2. There is a natural bijection between J(P ) and the set of antichains of P , since the maximal
elements of any order ideal form an antichain that generates it. (Recall that an antichain is a set of elements
that are pairwise incomparable.) Moreover, for each order ideal I, the order ideals covered by I in J(P ) are
18
precisely those of the form I ′ = I \ {x}, where x is a maximal element of I. In particular |I ′ | = |I| − 1 for all
such I ′ , and it follows by induction that J(P ) is ranked by cardinality.
We will shortly prove Birkhoff’s theorem (Theorem 1.3.7), a.k.a. the Fundamental Theorem of Finite Dis-
tributive Lattices: the finite distributive lattices are exactly the lattices of the form J(P ), where P is a finite
poset.
Definition 1.3.3. Let L be a lattice. An element x ∈ L is join-irreducible if it cannot be written as the join
of two other elements. That is, if x = y ∨ z then either x = y or x = z. The subposet (not sublattice!) of L
consisting of all join-irreducible elements is denoted Irr(L). Here is an example.
b f
e d b d
a c a c
L Irr(L)
If L is finite, then an element of L is join-irreducible if it covers exactly one other element. (This is not true
in a lattice such as R under the natural order, in which there are no covering relations!) The condition of
finiteness can be relaxed; see Problem 1.10.
Definition 1.3.4. A factorization of x ∈ L is an equation of the form
x = p1 ∨ · · · ∨ pn
In analogy with ring theory, call a lattice Artinian if it has no infinite descending chains. (For example, L
is Artinian if it is finite, or chain-finite, or locally finite and has a 0̂.) If L is Artinian, then every element
x ∈ L has a factorization — if x itself is not join-irreducible, express it as a join of two smaller elements,
then repeat. Moreover, every factorization can be reduced to an irredundant factorization by deleting each
factor strictly less than another (which does not change the join of the factors). Throughout the rest of the
section, we will assume that L is Artinian.
For general lattices, irredundant factorizations need not be unique. For example, the 1̂ element of M5 can
be factored irredundantly as the join of any two atoms. On the other hand, distributive lattices do exhibit
unique factorization, as we will soon prove (Proposition 1.3.6).
Proposition 1.3.5. Let L be a distributive lattice and let p ∈ Irr(L). Suppose that p ≤ q1 ∨ · · · ∨ qn . Then p ≤ qi
for some i.
Proof. By distributivity,
p = p ∧ (q1 ∨ · · · ∨ qn ) = (p ∧ q1 ) ∨ · · · ∨ (p ∧ qn )
and since p is join-irreducible, it must equal p ∧ qi for some i, whence p ≤ qi .
19
Proposition 1.3.5 is a lattice-theoretic analogue of the statement that if a prime p divides a product of positive
numbers, then it divides at least one of them. (This is in fact exactly what the result says when applied to
the divisor lattice Dn .)
Proposition 1.3.6 (Unique factorization for distributive lattices). Let L be a distributive lattice. Then every
x ∈ L can be written uniquely as an irredundant join of join-irreducible elements.
x = p1 ∨ · · · ∨ pn = q1 ∨ · · · ∨ qm (1.4)
Meanwhile, the join-irreducible order ideals in P are just the principal order ideals, i.e., those generated by
a single element. So the poset isomorphism P → Irr(J(P )) is given by
ψ(y) = ⟨y⟩.
These facts need to be checked; the details are left to the reader (Problem 1.12).
Corollary 1.3.8. Every finite distributive lattice L is graded.
Proof. The FTFDL says that L ∼= J(P ) for some finite poset P . Then L is ranked by Remark 1.3.2, and it is
bounded with 0̂ = ∅ and 1̂ = P .
Corollary 1.3.9. Let L be a finite distributive lattice. The following are equivalent:
1. L is a Boolean lattice.
2. Irr(L) is an antichain.
3. L is atomic (i.e., every element in L is the join of atoms). Equivalently, every join-irreducible element is an
atom.
4. L is complemented. That is, for each x ∈ L, there exists a unique element x̄ ∈ L such that x ∨ x̄ = 1̂ and
x ∧ x̄ = 0̂.
5. L is relatively complemented. That is, for every interval [y, z] ⊆ L and every x ∈ [y, z], there exists a unique
element u ∈ [y, z] such that x ∨ u = z and x ∧ u = y.
(4) =⇒ (3): Suppose that L is complemented, and suppose that y ∈ Irr(L) is not an atom. Let x be an atom
in [0̂, y]. Then
(x ∨ x̄) ∧ y = 1̂ ∧ y = y
(x ∨ x̄) ∧ y = (x ∧ y) ∨ (x̄ ∧ y) = x ∨ (x̄ ∧ y)
20
by distributivity. So y = x ∨ (x̄ ∧ y), which is a factorization of y, but y is join-irreducible, which implies
x̄ ∧ y = y, i.e., x̄ ≥ y. But then x̄ ≥ x and x̄ ∧ x = x ̸= 0̂, a contradiction.
(3) =⇒ (2): This follows from the observation that no two atoms are comparable.
Join and meet could have been interchanged throughout this section. For example, the dual of Proposi-
tion 1.3.6 says that every element in a distributive lattice L has a unique “cofactorization” as an irredundant
meet of meet-irreducible elements, and L is Boolean iff every element is the meet of coatoms. (In this case
we would require L to be Noetherian instead of Artinian — i.e., to contain no infinite increasing chains. For
example, Young’s lattice is Artinian but not Noetherian.)
Definition 1.4.1. A lattice L is modular if every x, y, z ∈ L with x ≤ z satisfy the modular equation:
x ∨ (y ∧ z) = (x ∨ y) ∧ z. (1.5)
Note that for all lattices, if x ≤ z, then x ∨ (y ∧ z) ≤ (x ∨ y) ∧ z. Modularity says that, in fact, equality holds.
x∨y z x∨y z
x y∧z x y∧z
Modular Non-modular
The term “modularity” arises in algebra: a canonical example of a modular lattice is the poset of modules
over any ring, ordered by inclusion (Corollary 1.4.3).
x ∨ (y ∧ z) = (x ∨ y) ∧ (x ∨ z) = (x ∨ y) ∧ z.
3. The lattice L is modular if and only if its dual L∗ is modular. Unlike the corresponding statement for
distributivity, this is immediate, because the modular equation is invariant under dualization.
4. The nonranked lattice N5 is not modular.
21
z
y
x
Here x ≤ z, but
x ∨ (y ∧ z) = x ∨ 0̂ = x,
(x ∨ y) ∧ z = 1̂ ∧ z = z.
In fact, N5 is the unique obstruction to modularity, as we will soon see (Thm. 1.4.5).
5. The nondistributive lattice M5 ∼= Π3 is modular. However, Π4 is not modular (exercise).
Theorem 1.4.2. [Characterizations of modularity] Let L be a lattice. Then the following are equivalent:
1. L is modular.
2. For all x, y, z ∈ L, if x ∈ [y ∧ z, z], then x = (x ∨ y) ∧ z.
3. For all x, y, z ∈ L, if x ∈ [y, y ∨ z], then x = (x ∧ z) ∨ y.
4. For all y, z ∈ L, the lattices L′ = [y ∧ z, z] and L′′ = [y, y ∨ z] are isomorphic, via the maps
α : L′ → L′′ β : L′′ → L′
q 7→ q ∨ y, p 7→ p ∧ z.
Proof. (1) =⇒ (2: If y ∧z ≤ x ≤ z, then the modular equation x∨(y ∧z) = (x∨y)∧z reduces to x = (x∨y)∧z.
b ∧ c ≤ a ∨ (b ∧ c) ≤ c ∨ c = c
(2) ⇐⇒ (3): These two conditions are duals of each other (i.e., L satisfies (2) iff L∗ satisfies (3)), and modu-
larity is a self-dual condition.
(2)+(3) ⇐⇒ (4): The functions α and β are always order-preserving functions with the stated domains and
ranges. Conditions (2) and (3) say respectively that β ◦ α and α ◦ β are the identities on L′ and L′′ ; together,
these conditions are equivalent to condition (4).
Corollary 1.4.3. Let R be a (not necessarily commutative) ring and M a (left) R-submodule. Then the (possibly
infinite) poset L(M ) of (left) R-submodules of M , ordered by inclusion, is a modular lattice with operations Y ∨ Z =
Y + Z and Y ∧ Z = Y ∩ Z.
[Y ∩ Z, Z] ∼
= L(Z/(Y ∩ Z)) ∼
= L((Y + Z)/Y ) ∼
= [Y, Y + Z]
22
In particular, the subspace lattices Ln (q) are modular (see Example 1.2.5).
Example 1.4.4. For a (finite) group G, let L(G) denote the lattice of subgroups of G, with operations H ∧K =
H ∩ K and H ∨ K = HK (i.e., the group generated by H ∪ K). If G is abelian then L(G) is always modular,
but if G is non-abelian then modularity can fail.
For example, let G = S4 , let X and Y be the cyclic subgroups generated by the cycles (1 2 3) and (3 4)
respectively, and let Z = A4 (the alternating group). Then (XY ) ∩ Z = Z but X(Y ∩ Z) = Z. Indeed, these
groups generate a sublattice of L(S4 ) isomorphic to N5 :
S4
A4
⟨(3 4)⟩
⟨(1 2 3)⟩
{Id}
Proof. Both =⇒ directions are easy, because distributivity and modularity are conditions inherited by
sublattices, and N5 is not modular and M5 is not distributive.
Suppose that x, y, z is a triple for which modularity fails. One can check that
x∨y
(x ∨ y) ∧ z
x∧y
Suppose that L is not distributive. If it isn’t modular then it contains an N5 , so there is nothing to prove. If
it is modular, then choose x, y, z such that
x ∧ (y ∨ z) > (x ∧ y) ∨ (x ∧ z).
23
1. this inequality is invariant under permuting x, y, z;
2. (x∧(y ∨z))∨(y ∧z) and the two other lattice elements obtained by permuting x, y, z form an antichain;
3. x ∨ y = x ∨ z = y ∨ z, and likewise for meets.
x∨y∨z
x∧y∧z
A corollary is that every modular lattice is graded, because a non-graded lattice must contain a sublattice
isomorphic to N5 . The details are left to the reader; we will eventually prove the stronger statement that
every semimodular lattice is graded.
Recall that the notation x ⋖ y means that x is covered by y, i.e., x < y and there exists no z strictly between
x, y (i.e., such that x < z < y).
x ∧ y ⋖ y =⇒ x ⋖ x ∨ y. (1.6)
Note that both upper and lower semimodularity are inherited by sublattices, and that L is upper semimod-
ular if and only if its dual L∗ is lower semimodular. Also, the implication (1.6) is trivially true if x and y
are comparable. If they are incomparable (as we will often assume), then there are several useful colloquial
rephrasings of semimodularity:
• “If meeting with x merely nudges y down, then joining with y merely nudges x up.”
• In the interval [x ∧ y, x ∨ y] ⊆ L pictured below, if the southeast relation is a cover, then so is the
northwest relation.
x∨y
•
x y (1.7)
⇐
=
•
x∧y
• This condition is often used symmetrically: if x, y are incomparable and they both cover x ∧ y, then
they are both covered by x ∨ y.
• Contrapositively, “If there is other stuff between x and x ∨ y, then there is also other stuff between
x ∧ y and y.”
24
Example 1.5.2. The partition lattice Πn is an important example of an upper semimodular lattice. To see
that it is USM, let π and σ be incomparable set partitions of [n], and suppose that σ ⋗ σ ∧ π. Recall that this
means that σ ∧ π can be obtained from σ by splitting some block B ∈ σ into two sub-blocks B ′ , B ′′ . More
specifically, we can write σ = A1 | · · · |Ak |B and σ ∧ π = A1 | · · · |Ak |B ′ |B ′′ , where B is the disjoint union of
B ′ and B ′′ . Since σ ∧ π refines π but σ does not, we know that A1 , . . . , Ak , B ′ , B ′′ are all subsets of blocks of
π but B is not; in particular B ′ and B ′′ are subsets of different blocks of π, say C ′ and C ′′ respectively. But
then merging C ′ and C ′′ produces a partition τ that covers π and is refined by σ, so it must be the case that
τ = σ ∨ π, and we have proved that Πn is USM. ◀
Lemma 1.5.3. If a lattice L is modular, then it is both upper and lower semimodular.
Proof. If x ∧ y ⋖ y, then the sublattice [x ∧ y, y] has only two elements. If L is modular, then condition (4) of
the characterization of modularity (Theorem 1.4.2) implies that [x ∧ y, y] ∼ = [x, x ∨ y], so x ⋖ x ∨ y. Hence L
is upper semimodular. The dual argument proves that L is lower semimodular.
In fact, upper and lower semimodularity together imply modularity. We will show that any of these three
conditions on a lattice L implies that it is graded, and that its rank function r satisfies
In other words, if it only takes one step to walk up from q to r, then it takes at most one step to walk from
q ∨ s to r ∨ s.
• If p = r, then q ∨ s ≥ r. So q ∨ s = r ∨ (q ∨ s) = (r ∨ q) ∨ s = r ∨ s.
• If p = q, then p = (q ∨ s) ∧ r = q ⋖ r. Applying semimodularity to the diamond figure below, we
obtain (q ∨ s) ⋖ (q ∨ s) ∨ r = r ∨ s.
r∨s
•
q∨s r
•
p = (q ∨ s) ∧ r = q
Theorem 1.5.5. Let L be a finite lattice. Then L is USM if and only if it is ranked, with rank function r satisfying
the submodular inequality or semimodular inequality
Proof. ( ⇐= ) Suppose that L is a ranked lattice with rank function r satisfying (1.8). Suppose that x, y are
incomparable and x∧y⋖y so that r(y) = r(x∧y)+1. Incomparability implies x∨y > x, so r(x∨y)−r(x) > 0.
On the other hand, rearranging (1.8) gives
25
1̂
xn−1 ym−1
L0 L00
x3 z3
x2 z2 y2
x1 y1
0̂
Let L′ = [x1 , 1̂] and L′′ = [y1 , 1̂]. (See Figure 1.6.) By induction, these sublattices are both ranked. Moreover,
c(L′ ) = n − 1. If x1 = y1 then Y and X are both saturated chains in the ranked lattice L′ and we are done,
so suppose that x1 ̸= y1 . Let z2 = x1 ∨ y1 . By (1.9), z2 covers both x1 and y1 . Let z2 , . . . , 1̂ be a saturated
chain in L (thus, in L′ ∩ L′′ ).
Since L′ is ranked and z ⋗ x1 , the chain z1 , . . . , 1̂ has length n − 2. So the chain y1 , z1 , . . . , 1̂ has length n − 1.
On the other hand, L′′ is ranked and y1 , y2 , . . . , 1̂ is a saturated chain, so it also has length n − 1. Therefore
the chain 0̂, y1 , . . . , 1̂ has length n as desired.
Second, we show that the rank function r of L satisfies (1.8). Let x, y ∈ L and take a saturated chain
x ∧ y = c0 ⋖ c1 ⋖ · · · ⋖ cn−1 ⋖ cn = x.
7 Recall that the length of a saturated chain is the number of minimal relations in it, which is one less than its cardinality as a subset
26
Note that n = r(x) − r(x ∧ y). Then there is a chain
y = c0 ∨ y ≤ c1 ∨ y ≤ · · · ≤ cn ∨ y = x ∨ y.
By Lemma 1.5.4, each ≤ in this chain is either an equality or a covering relation. Therefore, the distinct
elements ci ∨ y form a saturated chain from y to x ∨ y, whose length must be ≤ n. Hence
The same argument shows that L is lower semimodular if and only if it is ranked, with a rank function
satisfying the reverse inequality of (1.8).
Theorem 1.5.6. L is modular if and only if it is ranked, with rank function r satisfying the modular equality
Proof. If L is modular, then it is both upper and lower semimodular, so the conclusion follows by Theo-
rem 1.5.5. On the other hand, suppose that L is a lattice whose rank function r satisfies (1.10). Let x ≤ z ∈ L.
We already know that x ∨ (y ∧ z) ≤ (x ∨ y) ∧ z, so it suffices to show that these two elements have the same
rank. Indeed,
and
The following construction gives the prototype of a geometric lattice. Let k be a field, let V be a vector space
over k, and let E be a finite subset of V (with repeated elements allowed). We may as well assume that E
spans V , so in particular dim V < ∞. Say that a flat is a subset of E of the form W ∩ E, where W ⊆ E is a
vector subspace. Define the vector lattice of E as
L(E) ∼
= {kA : A ⊆ E}. (1.12)
(where kA denotes the vector subspace of V generated by A), via the map sending A 7→ kA. (Note that
different subspaces of W can have the same intersection with E, and different subsets of E can span the
same vector space.) The poset L(E) is easily checked to be a lattice under the operations
(W ∩ E) ∧ (X ∩ E) = (W ∩ X) ∩ E, (W ∩ E) ∨ (X ∩ E) = (W + X) ∩ E.
27
The elements of L(E) are called flats. Certainly E = V ∩ E is a flat, hence the top element of L(E). The
bottom element is O ∩ E, where O ⊆ V is the zero subspace; thus O ∩ E consists of the copies of the zero
vector in E.
The tricky thing about the isomorphism (1.12) is that it is not so obvious which elements of E are flats. For
every A ⊆ E, there is a unique minimal flat containing A, namely Ā := kA ∩ E — that is, the set of elements
of E in the linear span of A. On the other hand, if v, w, x ∈ E with v + w = x, then {v, w} is not a flat,
because any vector subspace that contains both v and w must also contain x. So, an equivalent definition of
“flat” is that A ⊆ E is a flat if no vector in E \ A is in the linear span of the vectors in A.
The lattice L(E) is ranked, with rank function r(A) = dim kA. It is upper semimodular (Problem 1.17) but is
not in general modular (see Example 1.6.3 below). On the other hand, L(E) is always an atomic lattice: every
element is the join of atoms. This is a consequence of the simple fact that k⟨v1 , . . . , vk ⟩ = kv1 + · · · + kvk .
This motivates the following definition:
Definition 1.6.1. A lattice L is geometric if it is (upper) semimodular and atomic. If L ∼
= L(E) for some set
of vectors E, we say that E is a (linear) representation of L.
For example, the set E = {(0, 1), (1, 0), (1, 1)} ⊆ F22 is a linear representation of the geometric lattice M5 .
(For that matter, so is any set of three nonzero vectors in a two-dimensional space over any field, provided
none is a scalar multiple of another.)
(An affine subspace of V is a translate of a vector subspace; for example, a line or plane not necessarily
containing the origin.) In fact, any lattice of the form Laff (E) can be expressed in the form L(Ê), where
Ê is a certain point set constructed from E (homework problem). However, the dimension of the affine
span of a set A ⊆ E is one less than its rank — which means that we can draw geometric lattices of rank 3
conveniently as planar point configurations. If L ∼= Laff (E), we could say that E is a (affine) representation
of L.
Example 1.6.2. Let E = {a, b, c, d}, where a, b, c are collinear but no other set of three points is. Then Laff (E)
is the lattice shown below (which happens to be modular).
abcd
a
d
abc ad bd cd
b
a b c d
c
Example 1.6.3. If E is the point configuration on the left with the only collinear triples {a, b, c} and {a, d, e},
then Laff (E) is the lattice on the right.
28
e abcde
abc bd be cd ce ade
d
b c a d e
∅
a b c
This lattice is not modular: consider the two elements bd and ce. ◀
Example 1.6.4. Recall from Example 1.5.2 that the partition lattice Πn is USM for all n. In fact it is geometric.
To see that it is atomic, observe that the atoms are the set partitions with n − 1 blocks, necessarily one
doubleton block and n − 2 singletons; let πij denote the atom whose doubleton block is {i, j}. Then every
set partition σ is the join of the set {πij : i ∼σ j}.
In fact, Πn is a vector lattice. Let k be any field, llet {e1 , . . . , en } be the standard basis of V = kn , let
pij = ei − ej for all 1 ≤ i < j ≤ n, and let E be the set of all such vectors pij .. Then in fact Πn ∼= L(E). The
atoms πij of Πn correspond to the atoms k⟨pij ⟩ of L(E); the rest of the isomorphism is left as Problem 1.18.
Note that this construction works over any field k.
More generally, if G is any simple graph on vertex set [n] then the connectivity lattice K(G) is isomorphic
to L(EG ), where EG = {aij : ij is an edge of G}. ◀
1.7 Exercises
Posets
Problem 1.1. (a) Prove that every nonempty interval in a Boolean lattice is itself isomorphic to a Boolean
lattice.
(b) Prove that every interval in the subspace lattice Ln (q) is isomorphic to a subspace lattice.
(c) Prove that every interval in the partition lattice Πn is isomorphic to a product of partition lattices.
(The product of posets P1 , . . . , Pk is the Cartesian product P1 × · · · × Pk , equipped with the partial
order (x1 , . . . , xk ) ≤ (y1 , . . . , yk ) if xi ≤Pi yi for all i ∈ [k].)
Solution: (a) Let [A, B] be an interval in Booln and define f (X) = X \ A for X ∈ [A, B]. I claim that f is
a poset isomorphism [A, B] → Bool[B\A] . It is evidently order-preserving, and it is a bijection because the
map g(Y ) = Y ∪ A is a two-way inverse.
(b) Let [V, W ] be an interval in Ln (q) and define f (X) = X/W for X ∈ [V, W ]. I claim that f is a poset
isomorphism [V, W ] → Lk (q), where k = dim W = dim V . It is evidently order-preserving, and it is a
bijection because the map g(Y ) = Y ⊕ A is a two-way inverse.
(c) Let [π, σ] be an interval in Πn . Let B1 , . . . , Bk be the blocks of σ; then π is obtained by splitting each
block of σ into one or more subblocks, so we can write
π = {B1,1 , . . . , B1,a1 , ..., Bk,1 , . . . , Bk,ak }
where each ai ≥ 1 and Bi = {Bi,1 , . . . , Bi,ai } is a set partition of Bi for all i ∈ [k].
29
For each ρ ∈ [π, σ], let ρi be the set of blocks of ρ contained in Bi . Then ρi can be regarded as a partition of
the set Bi (since ρi is obtained by merging some blocks in Bi together), giving a map
Solution: L (P + Q) is the set of shuffles of L (P ) with L (Q). (Let p = |P | and q = |Q|. Take the numbers in
[p + q] and paint p of them red and q of them blue. Then fill the red numbers withh a linear extension of P
and the blue numbers with a linear extension of Q.) In particular
p+q
|L (P + Q)| = |L (P )| · |L (Q)|.
p
Problem 1.3. Let n be a positive integer. Let Dn be the set of all positive-integer divisors of n (including n
itself), partially ordered by divisibility.
(a) Prove that Dn is a ranked poset, and describe the rank function.
(b) For which values of n is Dn (i) a chain; (ii) a Boolean lattice? For which values of n, m is it the case
that Dn ∼
= Dm ?
(c) Prove that Dn is a distributive lattice. Describe its meet and join operations and its join-irreducible
elements.
(d) Prove that Dn is self-dual, i.e., there is a bijection f : Dn → Dn such that f (x) ≤ f (y) if and only if
x ≥ y.
Qs
Solution: Let the prime factorization of n be i=1 pai i , where the pi are distinct primes and ai > 0. So the
Qs
elements of Dn are the numbers i=1 pbi i , where 0 ≤ bi ≤ ai for all i.
(a) The covering relations in Dn are x⋖y if y/x is prime. This implies that Dn is ranked, with rank function r
given by
s s
!
Y X
r pbi i = bi .
i=1 i=1
(b) The lattice Dn is a chain iff n is a prime power (i.e., s = 1). If n = pa then Dn is a chain of length a, while
for the converse, any two distinct primes dividing n are incomparable in Dn .
The lattice Dn is a Boolean lattice iff n is squarefree, i.e., ai = 1 for all i. In this case, the isomorphism
f : Dn → Boolr is given by
s
!
Y
f pbi i = {i ∈ [s] : bi = 1}.
i=1
On the other hand, if ai > 1 for some i then p2i ∈ Dn cannot be expressed as the join of atoms, so Dn is not
atomic and hence not Boolean.
30
Qs Qs
(c) For x, y ∈ Dn , let x = i=1 pbi i , y = i=1 pci i . The lattice operations on Dn are given by
s s
max(bi ,ci ) min(bi ,ci )
Y Y
x ∨ y = lcm(x, y) = pi and x ∧ y = gcd(x, y) = pi .
i=1 i=1
Moreover, Dn is distributive because min and max are distributive over each other: min(b, max(c, d)) =
max(min(b, c), min(b, c)) for all b, c, d in R (or indeed in any chain), and the identity remains true upon
switching min and max.
The join-irreducible elements of Dn are the prime powers, i.e., elements of the form pbi i with 0 < bi ≤ ai .
The only number covered by such a number is pibi −1 . In general, the number of elements covered by x
equals the number of primes dividing x.
(d) The function f (x) = n/x is the desired anti-automorphism. It is its own inverse, hence a bijection, and
x|y ⇐⇒ n/y|n/x.
Problem 1.4. Let G be a graph on vertex set V = [n]. Recall from Example 1.2.3 that the connectivity lattice
of a graph is the subposet K(G) of Πn consisting of set partitions in which every block induces a connected
subgraph of G. Prove that K(G) is a lattice. Is it a sublattice of Πn ?
Solution: First, K(G) is bounded: the 0̂ element is 1|2| · · · |n, while the 1̂ is the set partition whose blocks are
the vertex sets of the connected components of G. (So 1̂K(G) = 1̂Πn if and only if G is connected.) Moreover,
if X ⊆ V and Y ⊆ V induce connected subgraphs of G and X ∩ Y ̸= ∅, then so does X ∪ Y , which says
that the join operation in Πn preserves membership in K(G). So K(G) is a bounded join-semilattice and
therefore a lattice by Proposition 1.2.9. It is in general not a sublattice of Πn , however. For example, let G
be the 4-cycle graph, with vertices labeled 1, 2, 3, 4 in cyclic order, and let π = 123|4 and σ = 134|2. Both
these partitions belong to K(G), but their meet in Πn is 13|2|4, which does not. (In fact the meet in K(G) is
0̂.) An interesting corollary of this proof is that a sub-join-semilattice of a lattice need not be a sublattice!
Problem 1.5. Let A be a finite family of sets. For A′ ⊆ A, define ∪A′ = A∈A′ A. Let U (A) = {∪A′ : A′ ⊆
S
A}, considered as a poset ordered by inclusion.
(a) Prove that U (A) is a lattice. (Hint: Don’t try to specify the meet operation explicitly.)
(b) Construct a set family A such that U (A) is isomorphic to weak Bruhat order on S3 (see Example 2.11).
(c) Construct a set family A such that U (A) is not ranked.
(d) Is every finite lattice of this form?
Solution:
(a) U (A) is a bounded poset under set containment and it has a well-defined join (namely union), so it is
a lattice by Prop. 2.7.
(b) Take A = {ab, abc, bcd, cd}.
(c) Take A = {abc, ade, bde, cde}.
(d) Yes. For x ∈ L, define Ax = L \ [x, 1̂] and A = {Ax : x ∈ L}. Clearly x ≤ y iff Ax ⊆ Ay . Moreover,
A(x ∨ y) = L \ [x ∨ y, 1̂]
= L \ ([x, 1̂] ∩ [y, 1̂])
= (L \ [x, 1̂]) ∪ (L \ [y, 1̂])
= A(x) ∪ A(y).
Therefore A = U (A). Moreover, defining join as union makes U (A) into a join-semilattice, hence a
lattice with join given by union, and the map x 7→ Ax is a join-semilattice isomorphism L → U (A),
hence a lattice isomorphism.
31
Problem 1.6. For 1 ≤ i ≤ n − 1, let si be the transposition in Sn that swaps i with i + 1. (The si are
called elementary transpositions.) You probably know that {s1 , . . . , sn−1 } is a generating set for Sn (and if
you don’t, you will shortly prove it). For w ∈ Sn , an expression w = si1 · · · sik is called a reduced word if
there is no way to express w as a product of fewer than k generators.
(a) Show that every reduced word for w has length equal to inv(w) (as defined in (1.2)).
(b) Define a partial order ≺ on Sn as follows: w ≺ v if there exists a reduced word si1 · · · sik for v such
that w is the product of some proper subword w = sij1 · · · sijℓ . (Sorry about the triple subscripts; this
just means that v is obtained by deleting some of the letters from the reduced word for w.) Prove
that if w ≺ v, then w < v in Bruhat order. (The converse is true but requires significantly more work;
see [BB05], in particular Theorems 1.4.3 and 2.2.2.)
Solution: (a) Let ℓ(w) be the length of any (hence every) reduced word of w.
First, observe that w′ = si w is the word obtained by swapping the entries i and i + 1, wherever they appear
in w (written in one-line notation). The pair (i, i + 1) is an inversion of w if and only if it is not an inversion
of w′ ; on the other hand, every other pair is either an inversion of both w and w′ , or of neither of them (since
if j ∈ [n] \ {i, i + 1}, then j < i if and only if j < i + 1). Therefore, inv(wsi ) equals either inv(w) + 1 or
inv(w) − 1. In particular, if w can be expressed as the product of ℓ elementary transpositions, it can have at
most ℓ inversions, and we conclude that inv(w) ≤ ℓ(w).
Now we prove that inv(w) ≥ ℓ(w) by induction on k = inv(w). If k = 0 then w is the identity, which can be
expressed as the empty word, so ℓ(w) = 0. If w is not the identity, then there must be some pair (i, i + 1)
such that i + 1 precedes i in the one-line notation for w. (Otherwise, 1 must precede 2, which must precede
3, which. . . which means w = 12 · · · n.) By the previous observation, inv(si w) = k − 1, and by induction, si w
has a reduced word of length k − 1, say si w = sj1 · · · sjk−1 , which implies that w = si (si w) = si sj1 · · · sjk−1 .
This word might not be reduced, but it has length k, so a reduced word for w must have length ≤ k, i.e.,
ℓ(w) ≤ k = inv(w), which is the reverse equality. (So in fact that word for w was reduced after all.)
(b) For this part of the problem, we will eliminate a layer of subscripts by using si to mean “the ith letter in
a reduced word” rather than “the transposition that swaps i with i + 1”.
Suppose that s1 · · · sk is a reduced word for v. Let t = sk sk−1 · · · sj+1 sj sj+1 · · · sk−1 sk . Then t is a conjugate
of sj , hence a transposition. Letting w = vt, we have
by cancelling the underlined parts. So v = wt, and ℓ(w) < k = ℓ(v), so w < v in Bruhat order.
Problem 1.7. Prove that the rank-generating functions of weak order and Bruhat order on Sn are both
n
Y 1 − qi
.
i=1
1−q
(Hint: Induct on n, and use one-line notation for permutations, not cycle notation.)
Solution: Both weak order and Bruhat order are ranked, with rank function given by inversion number (1.2).
Therefore, we need to prove that
n
X Y 1 − qi
q inv(w) = .
i=1
1−q
w∈Sn
For n = 1, this equation reduces to 1 = 1. For n > 1, let g(w) be the permutation obtained from w by deleting
n; for example, if w = 1574362 ∈ S7 then g(w) = 154362 ∈ S6 . Then inv(g(w)) = inv(w) − (n − w−1 (n))
32
(since applying g destroys the inversions involving n and a later digit, but preserves all others). Moreover,
the map Sn → Sn−1 × [n] sending w to (g(w), w−1 (n)) is a bijection. Therefore, notating k = w−1 (n), we
see that
X n
X X
q inv(w) = q inv(v)+(n−k)
w∈Sn v∈Sn−1 k=1
X
= (q n−1 + q n−2 + · · · + q + 1)q inv(v)
v∈Sn−1
n−1
1 − qn Y 1 − qi
= (by induction)
1 − q i=1 1 − q
n
Y 1 − qi
= .
i=1
1−q
Distributive lattices
Problem 1.8. Prove that the two formulations (1.3a) and (1.3b) of distributivity of a lattice L are equivalent,
i.e.,
x ∧ (y ∨ z) = (x ∧ y) ∨ (x ∧ z) ∀x, y, z ∈ L ⇐⇒ x ∨ (y ∧ z) = (x ∨ y) ∧ (x ∨ z) ∀x, y, z ∈ L.
Solution: A distributive lattice L = J(P ) is isomorphic to a divisor lattice if and only if the poset P = Irr(L)
is a disjoint union of chains, i.e., has the form {xi,j : i ∈ [k], 0 ≤ j ≤ ak } for some positive integers a1 , . . . , ak ,
with xi,j ≤ xi′ ,j ′ iff i = i′ and j ≤ j ′ .
( =⇒ ): The join-irreducibles in Dn are the prime powers dividing n, and pi |q j if and only if p = q and i ≤ j.
33
Problem 1.10. Let L be a finite lattice and x ∈ L. Prove that x is join-irreducible if it covers exactly one
other element. What weaker conditions than “finite” suffice?
Solution: Suppose x covers both y and z. Then x is a minimal element of the set {a ∈ L : a ≥ y and a ≥ z},
and since L is a lattice it must be the minimal element, i.e., x = y ∨ z, which says precisely that x is not
join-irreducible.
W
Suppose that x only covers a single element y. Then every p < x satisfies p ≤ y, so A ≤ p for any
A ⊆ [0̂, x] \ {x}, which means that x cannot be expressed as a nontrivial join.
If x covers no other elements, then x = 0̂ is not join-irreducible. This is the part of the statement that
requires a finiteness assumption. Chain-finiteness suffices to conclude that any element that covers no
other elements is not join-irreducible. Local finiteness plus the existence of a 0̂ element is also sufficient.
Problem 1.11. Let Y be Young’s lattice (which we know is distributive).
Solution: (a) The join-irreducibles in Y are the rectangles, i.e., the partitions with only one part size, hence
only one corner. To see this, note that if a Ferrers diagram contains a particular box, then it contains all boxes
weakly northwest (i.e., neither south nor east) of it. If λ is a rectangle and λ = µ ∨ µ′ , then WLOG µ contains
the unique corner of λ, but then by the previous observation µ contains every box of λ, so µ ≥ λ ≥ µ and
equality holds.
(b) The Ferrers diagram of any partition λ can be covered by rectangular Ferrers diagrams whose southeast
corners are the corners of λ, and if λ is not itself a rectangle then each such rectangle is strictly smaller
than λ in Y . These rectangles are the maximal join-irreducibles in the irredundant factorization of λ, and k
is the number of southeast corners of λ. For example,
• = • ∨ ∨
• •
• •
is the unique irredundant factorization of λ = (5, 2, 2, 2, 1, 1) (the corners have been marked with colored
dots); note that there are k = 3 factors, which is the number of distinct parts of λ.
(c) Recall that a maximal chain in [∅, λ] can be represented by a standard tableau of shape λ. For the case
λ = (n, n), there is a bijection between such standard tableaux and Dyck paths of length 2n (paths from
(0, 0) to (2n, 0) consisting of n northeast and n southeast steps, never going below the x-axis): step k is
northeast or southeast according as the number k is in the top or bottom row. In particular, the number of
1 2n
maximal chains in (∅, (n, n)) is the nth Catalan number Cn = n+1 n .
(d) Again, it is convenient to think in terms of standard tableaux. In this case a tableau is completely
specified by choosing m numbers in the range [2, m + n + 1] to put in the “arm” (due east of the first box)
and placing the other n numbers in the range [2, m + n + 1] in the “leg” (due south). Therefore, the answer
is m+n
m .
34
Problem 1.12. Fill in the details in the proof of the FTFDL (Theorem 1.3.7) by showing the following facts.
(a) For a finite distributive lattice L, show that the map ϕ : L → J(Irr(L)) given by
ϕ(x) = ⟨p : p ∈ Irr(L), p ≤ x⟩
Solution: (a) First, note that the angle brackets could be replaced with set braces in the definition of ϕ. Since
x is the join of the elements in ϕ(x), it follows that ϕ is one-to-one. On the other hand, the generators of an
order ideal I ∈ Irr(L) form an antichain, so their join is an irredundant factorization of some element x ∈ L,
and then ϕ(x) = I. So ϕ is onto.
Join and meet in J(Irr(L)) are given by union and intersection respectively, and in addition, p ≤ x ∧ y if and
only if p ≤ x and p ≤ y, so
ϕ(x ∧ y) = {p ∈ Irr(L) : p ≤ x ∧ y}
= {p ∈ Irr(L) : p ≤ x, p ≤ y}
= {p ∈ Irr(L) : p ≤ x} ∩ {p ∈ Irr(L) : p ≤ x ≤ y}
= ϕ(x) ∧ ϕ(y).
Moreover, if p ≤ x or p ≤ y, then p ≤ x ∨ y, and the converse is true in a distributive lattice since p ∈ Irr(L),
by Proposition 1.3.5. Therefore, the previous calculation holds upon replacing ∧ with ∨ and ∩ with ∪. So ϕ
is an isomorphism of lattices.
(b) Let P be a poset and I ⊆ P an order ideal. If I = ⟨x⟩ = [∅, x], then I cannot be expressed irredundantly
as the join (that is, union) of smaller order ideals, because any ideal that contains x automatically contains I
as a subset. So I is join-irreducible in J(P ). On the other hand, if p, q are distinct maximal elements of I,
then I ′ = I \ {p} and I ′′ = I \ {q} are order ideals and proper subsets of I, and I = I ′ ∪ I ′′ , so I is not
join-irreducible.
Problem 1.13. Let L be a sublattice of Booln that is accessible: if S ∈ L\{∅} then there exists some x ∈ S such
that S \ {x} ∈ L. Construct a poset P on [n] such that J(P ) = L. (Notice that I wrote “= L”, not “∼ = L.” It is
not enough to invoke Birkhoff’s theorem to say that such a P must exist! The point is to explicitly construct
a poset P on [n] whose order ideals are the sets in L.)
Solution: The condition that L is a sublattice if Booln is precisely that it is closed under unions and intersec-
tions. Accordingly, for x ∈ [n], the set \
Ix = A
A∈L: x∈A
is a member of L, and it is the unique smallest element of L containing x (in particular it is nonempty).
Claim 1: If x ̸= y then Ix ̸= Iy .
— Since Ix is accessible, there is a set Ix′ ∈ L such that Ix′ = Ix \ {z} for some z. But since Ix is the smallest
set in L containing x, it must be the case that z = x. In particular, if Ix = Iy then x = y.
35
Claim 3: Every join-irreducible
S in L is of the form Ix for some x.
— Let A ∈ L. Then A = x∈A Ix , from which the claim follows.
Indeed, the ⊆ containment is clear. For ⊇, observe that if x ∈ O and y ∈ Ix then Iy ⊆ Ix (by minimality of
Iy ) so y ≤P x (by definition of P ) so y ∈ O (because O is an order ideal). We conclude that J(P ) is precisely
the sublattice of Booln generated by the Ix , namely L.
Modular lattices
Problem 1.14. Let Ln (q) be the poset of subspaces of an n-dimensional vector space over the finite field Fq
(so Ln (q) is a modular lattice by Corollary 1.4.3).
(a) Prove directly from the definition of modularity that Ln (q) is modular. (I.e., verify algebraically that
the join and meet operations obey the modular equation (1.5).)
(b) Prove the assertion in Example 1.2.5 that the number of k-dimensional subspaces of Fnq is nk q . Hint:
Every vector space of dimension k is determined by an ordered basis v1 , . . . , vk . How many ordered
bases does each k-dimensional vector space V ∈ Ln (q) have? How many sequences of vectors in Fnq
are ordered bases for some k-dimensional subspace?
(c) Count the maximal chains in Ln (q).
Solution: (a) Modularity is equivalent to the statement that if X, Y, Z are vector subspaces of V with X ⊆ Z,
then (X + Y ) ∩ Z = X + (Y ∩ Z). The ⊇ direction follows because both X and Y ∩ Z are subspaces of both
X + Y and Z. For the ⊆ direction, suppose that x ∈ X, y ∈ Y , and x + y ∈ Z. Since x ∈ Z, we must have
y ∈ Z as well. That is, y ∈ Y ∩ Z, so x + y ∈ X + (Y ∩ Z).
(b) The number of ordered bases for a k-dimensional space over Fq is (q k − 1)(q k − q) · · · (q k − q k−1 ),
because each vector is constrained not to lie in the space generated by its predecessors. For a similar reason,
the number of sequences of vectors that form an ordered basis for some k-dimensional space over Fq is
(q n − 1)(q n − q) · · · (q n − q k−1 ). The desired result follows.
(c) A maximal chain (also called a complete flag) is a sequence of subspaces 0 = V0 ⊂ V1 ⊂ · · · · Vn = Fnq
such that dim Vi = i for all i. Such a thing can be described by an ordered basis {v1 , . . . , vn } of V for which
Vi = Fq ⟨v1 , . . . , vi }, and we have counted these in part (b). On the other hand, the number of ordered bases
that correspond to any particular flag is (q − 1)(q 2 − q) · · · (q n − q n−1 ), where the ith factor is |Vi \ Vi−1 |. So
the answer is
(q n − 1)(q n − q) · · · (q n − q n−1 )
.
(q − 1)(q 2 − q) · · · (q n − q n−1 )
Solution: Recall that the rank function on Πn is r(π) = n − |π|. Let π = 12|34 and π ′ = 13|24. Then π ∨ π ′ =
1234 = 1̂ and π ∧ π ′ = 1|2|3|4 = 0̂. In particular r(π) + r(π ′ ) = 2 + 2 = 4 and r(π ∨ π ′ ) + r(π ∧ π ′ ) = 3 + 0 = 3.
36
Semimodular and geometric lattices
Problem 1.16. Let L be a lattice with the following property: for all x, y ∈ L, if x ∧ y is covered by both
x and y, then x ∨ y covers x and y. Prove that L is upper semimodular. (Obviously upper-semimodular
lattices have this property, so this exercise provides an alternative definition of upper semimodularity.)
Solution: Suppose that x, y are incomparable elements such that x ∧ y ⋖ y. We wish to show that x ⋖ x ∨ y.
Let C be a maximum chain in [x ∧ y, x], and let c > 0 be its length; we induct on c. If c = 1 then x ∧ y ⋗ x
and we are done by the assumption on L. Otherwise, let z be the element of C that covers x ∧ y, so that we
have the following picture:
x∨y
x z∨y
• •
z y
• •
x∧y =z∧y
By construction both z and y cover z ∧ y, so by assumption they are both covered by z ∨ y (notated by the
•s in the figure). In particular, (z ∨ y) ∧ x must equal z, and certainly (z ∨ y) ∨ x = x ∨ y. On the other hand,
C ′ = C \ {x ∧ y} is a maximum chain in [z, x ∨ y], and its length is strictly less than that of C, so by induction
x ∨ y covers x, as desired.
Problem 1.17. Prove that the lattice L(E) defined in (1.11) is upper semimodular.
Solution: Let A, B be flats, and let Y = kA and Z = kB. Recall that dim Y +dim Z = dim(Y +Z)+dim(Y ∩Z)
(this is the statement that subspace lattices are modular). Then Y + Z = k(A ∪ B) = k(A ∨ B) (since A ∨ B
is the smallest flat containing A ∪ B, namely k(A ∪ B) ∩ E). On the other hand, A ∩ B = A ∧ B ⊆ Y ∩ Z,
but equality may not hold. Therefore, the rank function r of L(E) satisfies
This space is certainly an element of L(E). Also, observe that Vπ ⊆ Vσ whenever π refines σ, because then Vπ
is spanned by a subset of a spanning set of Vσ . The second expression for Vπ shows that it is defined by |π|
independent linear equations, hence has dimension n − |π|, which is exactly its rank in Πn .
37
At this point we have a one-to-one map f : Πn → L(E) that preserves rank and the order relation. We just
need to show that it is surjective, i.e., to show that every space spanned by a subset of E is of the form Vπ
for some π. Observe that the orthogonal complement k⟨pij ⟩⊥ with respect to the standard inner product
on V is the hyperplane
Hij = {x = (x1 , . . . , xn ) ∈ V : xi = xj }.
Therefore, for any subset F ⊆ E, we have
r r r
!⊥
X X \
k⟨F ⟩ = Lik jk = Hi⊥k jk = Hik jk .
k=1 k=1 k=1
On the other hand, this intersection is just defined by a collection of equations of the form xi = xj . These
give rise to an equivalence relation on [n], which is the same thing as a set partition π, and we see that
k⟨F ⟩ = Vπ .
It is worth noting that the construction is exactly the same over all fields F.
Problem 1.19. The purpose of this exercise is to show that the constructions L and Laff produce the same
class of lattices. Let k be a field and let E = {e1 , . . . , en } ⊆ kd .
(a) The augmentation of a vector ei = (ei1 , . . . , eid ) is the vector ẽi = (1, ei1 , . . . , eid ) ∈ kd+1 . Prove that
Laff (E) = L(Ẽ), where Ẽ = {ẽ1 , . . . , ẽn }.
(b) Let v be a vector in kd that is not a scalar multiple of any ei , let H Let H ⊆ kd be a generic affine
hyperplane, let êi be the projection of ei onto H, and let Ê = {ê1 , . . . , ên }. Prove that L(E) = Laff (Ê).
(The first part is figuring out what “generic” means. A generic hyperplane might not exist for all
fields, but if k is infinite then almost all hyperplanes are generic.)
Solution: To be written
Problem 1.20. Recall from Corollary 1.3.9 that a lattice L is relatively complemented if, whenever y ∈ [x, z] ⊆
L, there exists u ∈ [x, z] such that y ∧ u = x and y ∨ u = z. Prove that a finite semimodular lattice is atomic
(hence geometric) if and only if it is relatively complemented.
(Here is the geometric interpretation of being relatively complemented. Suppose that V is a vector space,
L = L(E) for some point set E ⊆ V , and that X ⊆ Y ⊆ Z ⊆ V are vector subspaces spanned by flats of
L(E). For starters, consider the case that X = O. Then we can choose a basis B of the space Y and extend
it to a basis B ′ of Z, and the vector set B ′ \ B spans a subspace of Z that is complementary to Y . More
generally, if X is any subspace, we can choose a basis B for X, extend it to a basis B ′ of Y , and extend B ′
to a basis B ′′ of Z. Then B ∪ (B ′′ \ B ′ ) spans a subspace U ⊆ Z that is relatively complementary to Y , i.e.,
U ∩ Y = X and U + Y = Z.)
Solution: ( =⇒ ) Suppose that L is atomic. Let y ∈ [x, z], and choose u ∈ [x, z] such that y ∧ u = x (for
instance, u = x). If y ∨ u = z then we are done. Otherwise, choose an atom a ∈ L such that a ≤ z but
a ̸≤ y ∨ u. Set u′ = u ∨ a. By semimodularity u′ ⋗ u. Then u′ ∨ y ⋗ u ∨ y by Lemma 1.5.4, and u′ ∧ y = x (this
takes a little more work; see below). By repeatedly replacing u with u′ if necessary, we eventually obtain a
complement for y in [x, z].
Here is the verification that u′ ∧ y = x. First, by the submodular inequality, u′ ∧ y either equals or covers
u ∧ y = x. Suppose that u ∧ y ⋗ x. By atomicity, there is some atom b such that b ≤ u′ ∧ y but b ̸≤ u ∧ y. The
first relation implies that b ≤ u′ and b ≤ y; therefore, b ̸≤ u (else b ≤ u ∧ y). Now
u ⋖ u ∨ b ≤ u′ ∨ b = u′
38
(where the covering relation comes from semimodularity). But u′ ⋗ u, so the ≤ in this equation must be
equality, which says that
u′ ∨ y = u′ ∨ (b ∨ y) = (u′ ∨ b) ∨ y = u ∨ b ∨ y = u ∨ y
( ⇐= ) Suppose that L is relatively complemented and let x ∈ L. We want to write x as the join of atoms. If
x = 0̂ then it is the empty join; otherwise, let a1 ≤ x be an atom and let x1 be a complement for a1 in [0̂, x].
Then x1 < x and x = a1 ∨ x1 . Replace x with x1 and repeat, getting
39
Chapter 2
Poset Algebra
Throughout this chapter, every poset we consider will be assumed to be locally finite, i.e., every interval is
finite.
Let P be a poset and let Int(P ) denote the set of (nonempty) intervals of P . Recall that an interval is a subset
of P of the form [x, y] := {z ∈ P : x ≤ z ≤ y}; if x ̸≤ y then [x, y] = ∅.
Definition 2.1.1. The incidence algebra I(P ) is the set of functions α : Int(P ) → C (“incidence functions”)1 ,
made into a C-vector space with pointwise addition, subtraction and scalar multiplication. It is equivalent
to think of I(P ) as the set of functions on P × P , with α(x, y) = 0 if x ̸≤ y — this lets us write α(x, y) instead
of the more awkward α([x, y]). We make I(P ) into a ring with the convolution product:
X
(α ∗ β)(x, y) = α(x, z)β(z, y).
z∈[x,y]
Note that the assumption of local finiteness is both necessary and sufficient for convolution to be well-
defined for all incidence functions.
40
Proof. The basic idea is to reverse the order of summation:
X
[(α ∗ β) ∗ γ](x, y) = (α ∗ β)(x, z) · γ(z, y)
z∈[x,y]
X X
= α(x, w)β(w, z) γ(z, y)
z∈[x,y] w∈[x,z]
X
= α(x, w)β(w, z)γ(z, y)
w,z: x≤w≤z≤y
X X
= α(x, w) β(w, z)γ(z, y)
w∈[x,y] z∈[w,y]
X
= α(x, w) · (β ∗ γ)(w, y)
w∈[x,y]
= [α ∗ (β ∗ γ)](x, y).
The ring I(P ) has a multiplicative identity, namely the Kronecker delta function, regarded as an incidence
function: (
1 if x = y,
δ(x, y) =
0 if x ̸= y.
Therefore, we sometimes write 1 for δ.
Once you know a ring has a multiplicative identity, the next natural question is which elements are invert-
ible. This question has a nice answer:
Proposition 2.1.3. An incidence function α ∈ I(P ) has a left/right/two-sided convolution inverse if and only if
α(x, x) ̸= 0 for all x (the “nonzero condition”). In that case, the inverse is given by the recursive formula
α(x, x)−1
if x = y,
−1
α (x, y) = −1
P −1 (2.1)
−α(y, y)
α (x, z)α(z, y) if x < y.
z: x≤z<y
This formula is well-defined by induction on the size of [x, y], with the cases x = y and x ̸= y serving as the
base case and inductive step respectively.
Proof. Let β be a left convolution inverse of α. In particular, α(x, x) = β(x, x)−1 for all x (use the equation
(α ∗ β)(x, x) = δ(x, x) = 1), so the nonzero condition is necessary.
and solving for β(x, y) (by pulling the z = y term out of the sum) gives the formula (2.1), which is well-
defined provided that α(y, y) ̸= 0. So the nonzero condition is also sufficient.
A similar argument shows that the nonzero condition is necessary and sufficient for α to have a right
convolution inverse. Moreover, the left and right inverses coincide: if β ∗ α = δ = α ∗ γ then β = β ∗ δ =
β ∗ α ∗ γ = γ by associativity.
41
Now we have a ring in which algebraic identities can encode facts about the poset P . We need some
interesting incidence functions to play with. The zeta function and eta function of P are defined as
( (
1 if x ≤ y, 1 if x < y,
ζ(x, y) = η(x, y) =
0 if x ̸≤ y, 0 if x ̸< y,
These trivial-looking incidence functions are useful because their convolution powers count important
things, namely multichains and chains in P . In other words, enumerative questions about posets can be
expressed algebraically. Specifically,
X X
ζ 2 (x, y) = ζ(x, z)ζ(z, y) = 1
z∈[x,y] z∈[x,y]
= #{z : x ≤ z ≤ y},
X X X
3
ζ (x, y) = ζ(x, z)ζ(z, w)ζ(w, y) = 1
z∈[x,y] w∈[z,y] x≤z≤w≤y
= #{(z, w) : x ≤ z ≤ w ≤ y},
k
ζ (x, y) = #{(x1 , . . . , xk−1 ) : x ≤ x1 ≤ x2 ≤ · · · ≤ xk−1 ≤ y}.
That is, ζ k (x, y) counts the number of multichains of length k between x and y (chains with possible re-
peats). If we replace ζ with η, then the calculations all work the same way, except that all the ≤’s are
replaced with <’s, so we get
η k (x, y) = #{(x1 , . . . , xk−1 ) : x < x1 < x2 < · · · < xk−1 < y},
the number of chains of length k (not necessarily saturated) between x and y. In particular, if the chains of P
are bounded in length (e.g., if P is finite), then η n = 0 for n ≫ 0.
Direct products of posets play nicely with the incidence algebra construction. Specifically, let P, Q be
bounded finite posets. For α ∈ I(P ) and ϕ ∈ I(Q), define αϕ ∈ I(P × Q) by
This defines a linear transformation F : I(P ) ⊗ I(Q) → I(P × Q). 2 In other words, (α + β)ϕ = αϕ + βϕ,
and α(ϕ + ψ) = αϕ + αψ, and α(cϕ) = (cα)ϕ = c(αϕ) for all c ∈ C. It is actually a vector space isomorphism,
because there is a bijection Int(P ) × Int(Q) → Int(P × Q) given by (I, J) → I × J, and F (χI ⊗ χJ ) = χI×J
(where χI is the characteristic function of I, i.e., the incidence function that is 1 on I and zero on other
intervals). In fact, more is true:
Proposition 2.1.4. The map F just defined is a ring isomorphism. That is, for all α, β ∈ I(P ) and ϕ, ψ ∈ I(Q),
αϕ ∗ βψ = (α ∗ β)(ϕ ∗ ψ).
Furthermore, the incidence functions δ and ζ are multiplicative on direct products, i.e.,
δP ×Q = δP δQ and ζP ×Q = ζP ζQ .
2 See §8.5 for an extremely brief introduction to the tensor product operation ⊗.
42
Proof. Let (x, x′ ) and (y, y ′ ) be elements of P × Q. Then
X
(αϕ ∗ βψ)[(x, x′ ), (y, y ′ )] = αϕ[(x, x′ ), (z, z ′ )] · βψ[(z, z ′ ), (y, y ′ )]
(z,z ′ )∈[(x,x′ ),(y,y ′ )]
X X
= α(x, z)ϕ(x′ , z ′ )β(z, y)ψ(z ′ , y ′ )
z∈[x,y] z ′ ∈[x′ ,y ′ ]
X X
= α(x, z)β(z, y) ϕ(x′ , z ′ )ψ(z ′ , y ′ )
z∈[x,y] z ′ ∈[x′ ,y ′ ]
The Möbius function µP of a poset P is defined as the convolution inverse of its zeta function: µP = ζP−1 .
This turns out to be one of the most important incidence functions on a poset. For a bounded poset, we
abbreviate µP (x) = µP (0̂, x) and µ(P ) = µP (0̂, 1̂). Proposition 2.1.3 provides a recursive formula for µ:
0
if y ̸≥ x (i.e., if [x, y] = ∅),
µ(x, y) = 1 if y = x, (2.2)
P
− z: x≤z<y µ(x, z) if x < y.
This is equivalent to the familiar recursive formula: to find µP (x), add up the values of µP at all elements
< x, then change the sign.
Example 2.2.1. If P = {0 < 1 < 2 < · · · } is a chain, then its Möbius function is given by µ(x, x) = 1,
µ(x, x + 1) = −1, and µ(x, y) = 0 otherwise. ◀
Example 2.2.2. Here are the Möbius functions µP (x) = µP (0̂, x) for the lattices N5 and M5 :
1 2
0
−1 −1 −1 −1
−1
N5 M5
1 1
And here are the Boolean lattice Bool3 and the divisor lattice D24 :
43
0 24
−1
0 12 8 0
1 1 1
1 6 4 0
−1 −1 −1
−1 3 2 −1
1 Bool3 1 1 D24
◀
Example 2.2.3 (Möbius functions of partition lattices). What is µ(Πn ) in terms of n? Clearly µ(Π1 ) = 1 and
µ(Π2 ) = −1, and µ(Π3 ) = µ(M5 ) = 2. For n = 4, we calculate µ(Π4 ) from (2.2). The value of µΠ4 (0̂, π)
depends only on the block sizes of π, in fact, [0̂, π] ∼
= Ππ1 × · · · × Ππk . We will use the fact that the Möbius
function is multiplicative on direct products; we will prove this shortly (Prop. 2.2.5).
Adding up the last column and multiplying by −1 gives µ(Π5 ) = 24. At this point you might guess that
µ(Πn ) = (−1)n−1 (n − 1)!, and you would be right. We will prove this soon. ◀
The Möbius function is useful in many ways. It can be used to formulate a more general version of
inclusion-exclusion called Möbius inversion. It behaves nicely under poset operations such as product, and
has geometric and topological applications. Even just the single number µ(P ) = µP (0̂, 1̂) tells you a lot
about a bounded poset P . Confusingly, this number itself is sometimes called the “Möbius function” of P
(I prefer “Möbius number” to avoid ambiguity). Here is the reason.
Definition 2.2.4. A family F of posets is hereditary if, for each P ∈ F , every interval in P is isomorphic to
some [other] poset in F . It is semi-hereditary if every interval in a member of F is isomorphic to a product
of members of F .
For example, the families of Boolean lattices, divisor lattices, and subspace lattices are all hereditary, and
the family of partition lattices is semi-hereditary (Problem 1.1). Knowing the Möbius number for every
44
poset in a hereditary family is equivalent to knowing their full Möbius functions. The same is true for
semi-hereditary families, for the following reason.
Proposition 2.2.5. The Möbius function is multiplicative on direct products, i.e., µP ×Q = µP µQ (in the notation of
Proposition 2.1.4).
Proof.
ζP ×Q ∗ µP µQ = ζP ζQ ∗ µP µQ = (ζP ∗ µP )(ζQ ∗ µQ ) = δP δQ = δP ×Q
which says that µP µQ = ζP−1×Q = µP ×Q . Here the second equality is the definition of multiplication in a
tensor product of rings. (It is also possible to prove that µP µQ = µP ×Q directly from the definition; this is
Problem 2.2.)
Since µ(Bool1 ) = −1 and Booln is a product of n copies of Bool1 , an immediate consequence of Proposi-
tion 2.2.5 is the formula
µ(Booln ) = (−1)n .
This can also be proved by induction on n (with the cases n = 0 and n = 1 easy). If n > 0, then
n−1
X X n
µ(Booln ) = − (−1)|A| = − (−1)k (by induction)
k
A⊊[n] k=0
n
X n
= (−1)n − (−1)k
k
k=0
= (−1)n − (1 − 1)n = (−1)n .
In particular, the full Möbius function of the Boolean lattice BoolS is given by µ(A, B) = µ(Bool|B\A| ) =
(−1)|B\A| for all A ⊆ B ⊆ S.
Example 2.2.6. Let P be a product of k chains of lengths a1 , . . . , ak . Equivalently,
ordered by x ≤ y iff xi ≤ yi for all i. (Recall that the length of a chain is the number of covering relations,
which is one less than the number of elements; see Definition 1.1.6.) Then Prop. 2.2.5 together with the
formula for the Möbius function of a chain (above) gives
(
0 if xi ≥ 2 for at least one i;
µ(0̂, x) = s
(−1) if x consists of s 1’s and k − s 0’s.
(The Boolean lattice is the special case that ai = 1 for every i.) This conforms to the definition of Möbius
function that you may have seen in enumerative combinatorics or number theory, since products of chains
are precisely divisor lattices. As mentioned above, the family of divisor lattices is hereditary: [a, b] ∼
= Db/a
for all a, b ∈ Dn with a|b. ◀
Here are a couple of enumerative applications of the Möbius function. The first, known as Philip Hall’s
Theorem,3 makes the connection between the Möbius function and topology more explicit.
3 Not to be confused with the unrelated Hall’s Marriage Theorem.
45
Theorem 2.2.7 (Philip Hall’s Theorem). [Sta12, Prop. 3.8.5] Let P be a finite bounded poset with at least two
elements. For k ≥ 1, let
Proof. Recall that ck = η k (0̂, 1̂) = (ζ −δ)k (0̂, 1̂). The trick is to use the geometric series expansion 1/(1+h) =
1 − h + h2 − h3 + h4 − · · · . Clearing both denominators and replacing h with η and 1 with δ, we get
∞
!
X
(δ + η)−1 = (−1)k η k .
k=0
The RHS looks like an infinite power series, but it is actually a polynomial, because η k = 0 for k sufficiently
large. (here is where we need the assumption that P is finite). So we have a valid equation in I(P ) (which
you can verify by multiplying δ + η by the RHS). Switching the two sides and evaluating on [0̂, 1̂] gives
∞
X ∞
X
(−1)k ck = (−1)k η k (0̂, 1̂) = (δ + η)−1 (0̂, 1̂) = ζ −1 (0̂, 1̂) = µ(0̂, 1̂).
k=0 k=0
This alternating sum looks like an Euler characteristic (see (6.2) below). In fact it is.
Corollary 2.2.8. Let P be a finite bounded poset with at least two elements, and let ∆(P ) be its order complex,
i.e., the simplicial complex (see Example 1.1.11) whose vertices are the elements of P \ {0̂, 1̂} and whose simplices are
chains. Each chain x0 = 0̂ < x1 < · · · < xk = 1̂ gives rise to a simplex {x1 , . . . , xk−1 } of ∆(P ) of dimension k − 2.
Hence fk−2 (∆(P )) = ck (P ) for all k ≥ 1, and the reduced Euler characteristic of ∆(P ) is
def X X
χ̃(∆(P )) ≡ (−1)k fk (∆(P )) = (−1)k−2 ck (P ) = µP (0̂, 1̂).
k≥−1 k≥1
Example 2.2.9. For the Boolean lattice P = Bool3 (see Example 2.2.2), we have c0 = 0, c1 = 1, c2 = 6, c3 = 6,
and ck = 0 for k > 3. Indeed, c0 − c1 + c2 − c3 = −1 = µP (0̂, 1̂). ◀
Proof. This is immediate from Philip Hall’s Theorem, since ck (P ) = ck (P ∗ ) for all k. (One can also prove
this fact by comparing the algebras I(P ) and I(P ∗ ); see Problem 2.3.)
The following result is one of the most frequent applications of the Möbius function.
46
Theorem 2.3.1 (Möbius inversion formula). Let P be a locally finite4 poset, let V be any C-vector space (usually,
but not always, C itself) and let f, g : P → V . Then
X X
g(x) = f (y) ∀x ∈ P ⇐⇒ f (x) = µ(y, x)g(y) ∀x ∈ P, (2.3a)
y: y≤x y: y≤x
X X
g(x) = f (y) ∀x ∈ P ⇐⇒ f (x) = µ(x, y)g(y) ∀x ∈ P. (2.3b)
y: y≥x y: y≥x
Proof. Stanley calls the proof “A trivial observation in linear algebra”. Let V be the vector space of functions
f : P → C. Consider the right action • and the left action • of I(P ) on V by
X
(f • α)(x) = α(y, x)f (y),
y: y≤x
X
(α • f )(x) = α(x, y)f (y).
y: y≥x
In terms of these actions, formulas (2.3a) and (2.3b) are respectively just the “trivial” observations
g = f •ζ ⇐⇒ f = g • µ, (2.4a)
g = ζ •f ⇐⇒ f = µ • g. (2.4b)
We just have to prove that these colored dots indeed define actions, i.e.,
f • (α ∗ β) = (f • α) • β and (α ∗ β) • f = α • (β • f ).
which is nothing more or less than the inclusion-exclusion formula. So Möbius inversion can be thought of
as a generalized form of inclusion-exclusion in which the Boolean lattice is replaced by an arbitrary locally
finite poset P . If we know the Möbius function of P , then knowing a combinatorial formula for either f
or g allows us to write down a formula for the other one. This is frequently useful when we can express an
enumerative problem in terms of a function on a poset.
4 In fact (2.3a) requires only that every principal order ideal is finite (for (2.3b), every principal order filter).
47
Remark 2.3.2. The proof of Möbius inversion goes through more generally for functions f, g : P → X,
where X is any C-vector space (for example, polynomials over C).
Example 2.3.3. Here’s an oldie-but-goodie. A derangement is a permutation σ ∈ Sn with no fixed points.
If Dn is the set of derangements in Sn , then
|D1 | = 0,
|D2 | = 1 = |{21}|,
|D3 | = 2 = |{231, 312}|,
|D4 | = 9 = |{2341, 2314, 2413, 3142, 3412, 3421, 4123, 4312, 4321}|,
...
Thus Dn = f (∅).
It is easy to calculate g(S) directly. If s = |S|, then a permutation fixing the elements of S is equivalent to a
permutation on [n] \ S, so g(S) = (n − s)!.
Rewritten in the incidence algebra I(2[n] ), this is just g = ζ • f . Thus f = µ • g, or in terms of the Möbius
inversion formula (2.3b),
n
n−s
X X X
f (S) = µ(S, R)g(R) = (−1)|R|−|S| (n − |R|)! = (−1)r−s (n − r)! .
r=s
r−s
R⊇S R⊇S
The number of derangements is then f (∅), which is given by the well-known formula
n
X n
(−1)r (n − r)!
r=0
r
◀
Example 2.3.4. As a number-theoretic application, we will use Möbius inversion to compute the closed
formula for Euler’s totient function
Let n = pa1 1 · · · pas s be the prime factorization of n, and let P = {p1 , . . . , ps }. We work in the lattice Dn ∼
=
Ca1 × · · · × Cas . Warning: To avoid confusion with the cardinality symbol, we will use the symbol ≤ to
mean the order relation in Dn : i.e., x ≤ y means that x divides y. For x ∈ Dn , define
48
Applying formulation (2.3b) of Möbius inversion gives
X
f (x) = µ(x, y)g(y).
y≥x
On the other hand g(x) = n/x, since x ≤ gcd(a, n) iff a is a multiple of x. Moreover, ϕ(n) = f (1), and
(
(−1)q if y is a product of distinct elements of P ,
µ(1, y) =
0 otherwise (i.e., if p2i ≤ y for some i).
Therefore,
X
ϕ(n) = f (1) = µ(1, y)(n/y)
y∈Dn
X (−1)|Q|
=n Q
Q⊆P pi ∈Q pi
n X p1 · · · pr
= (−1)|Q| Q
p1 · · · pr pi ∈Q pi
Q⊆P
n X Y
= (−1)r−|S| pi
p1 · · · pr
S=P \Q⊆P pi ∈S
r
n Y
= (−1)r (1 − pi )
p1 · · · pr i=1
= pa1 1 −1 · · · par r −1 (p1 − 1) · · · (pr − 1)
as is well known. ◀
Example 2.3.5. Let G = (V, E) be a finite graph with V = [n]. We may as well assume that G is simple (no
loops or parallel edges) and connected. A coloring of G with t colors, or for short a t-coloring, is just a
function κ : V (G) → [t]. An edge xy is monochromatic with respect to κ if κ(x) = κ(y), and a coloring is
proper if it has no monochromatic edges. What can we say about the number pG (t) of proper t-colorings?
This question can be expressed in terms of the connectivity lattice K(G) (see Example 1.2.3 and Problem 1.4).
For each t-coloring κ, let Gκ be the subgraph of G induced by the monochromatic edges, and let P (κ) be
the set partition of V (G) whose blocks are the components of Gκ ; then P (κ) is an element of K(G). The
coloring κ is proper if and only if P (κ) = 0̂K(G) , the partition of V (G) into singleton blocks. Accordingly, if
we define f : K(G) → N≥0 by
then the number of proper t-colorings is f (0̂). We can find another expression for this number by Möbius
inversion. Let X
g(π) = #{κ : P (κ) ≥ π} = f (σ).
σ≥π
The condition P (κ) ≥ π is equivalent to saying that the vertices in each block of π are colored the same. The
number of such colorings is just t|π| (choosing a color for each block, not necessarily different). Therefore,
Möbius inversion (version (2.3b)) says that
X X
pG (t) = f (0̂) = µ(0̂, π)g(π) = µ(0̂, π)t|π| . (2.5)
π∈K(G) π∈K(G)
49
While this formula is not necessarily easy to calculate, it does show that pG (t) is a polynomial in t; it is
called the chromatic polynomial. (There are other ways to show this fact.)
If G = Kn is the complete graph, then the connectivity lattice K(Kn ) is just the full partition lattice Πn . On
the other hand, we can calculate the chromatic polynomial of Kn directly: it is pKn (t) = t(t−1)(t−2) · · · (t−
n + 1) (since a proper coloring must assign different colors to all vertices). Combining this observation
with (2.5) gives X
µ(0̂, π)t|π| = t(t − 1)(t − 2) · · · (t − n + 1).
π∈K(Kn )
This is an identity of polynomials in t. Extracting the coefficients of the lowest degree (t1 ) terms on each
side gives
µ(0̂, 1̂) = (−1)n−1 (n − 1)!
so we have calculated the Möbius number of the partition lattice! There are many other ways to obtain this
result. ◀
Example 2.3.6. Here is another way to use Möbius inversion to compute the Möbius function itself. In this
example, we will do this for the lattice Ln (q).
For small n, it is possible to work out the Möbius function of Ln (q) by hand. For instance, µ(L1 (q)) =
µ(Bool1 ) = −1, and L2 (q) is a poset of rank 2 with q + 1 elements in the middle (since each line in F2q
is defined by a nonzero vector up to scalar multiples, so there are (q 2 − 1)/(q − 1) lines), so µ(L2 (q)) =
−(−(q+1)+1) = q. With a moderate amount of effort, one can check that µ(L3 (q)) = −q 3 and µ(L4 (q)) = q 6 .
Here is a way to calculate µ(Ln (q)) for general n, which will lead into the discussion of the characteristic
polynomial of a ranked poset.
Let V = Fnq , let L = Ln (q) (ranked by dimension) and let X be a Fq -vector space of cardinality t (yes,
cardinality, not dimension!) Let
so that X
g(W ) = f (U )
U ⊇W
For this last count, choose an ordered basis {v1 , . . . , vn } for V , and send each vi to a vector in X not in the
linear span of {ϕ(v1 ), . . . , ϕ(vi−1 )}; there are t − q i−1 such vectors. The identity (2.6) holds for infinitely
50
many integer values of t and is thus an identity of polynomials in the ring Q[t]. Therefore, it remains true
upon setting t to 0 (even though no vector space can have cardinality zero!), whereupon the second and
fourth terms in the equality (2.6) become
n
µLn (q) (0̂, 1̂) = (−1)n q ( 2 )
which is consistent with the n ≤ 4 cases given at the start of the example. ◀
The two previous examples suggest that in order to understand a finite graded poset P , one should study
the following polynomial.
Definition 2.3.7. Let P be a graded poset with rank function r. Its characteristic polynomial is
X
χ(P ; t) = µ(0̂, x)tr(1̂)−r(x) .
x∈P
In particular,
χ(P, 0) = µ(P ). (2.7)
In fact, since the Möbius function is multiplicative on direct products of posets (Proposition 2.2.5), so is the
characteristic polynomial.
The characteristic polynomial generalizes the Möbius number of a poset and contains additional infor-
mation as well. For example, let A be a hyperplane arrangement in Rn : a finite collection of affine linear
spaces of dimension n − 1. The arrangement separates Rn into regions, the connected components of
X = R \ H∈A H. Let P be the poset of intersections of hyperplanes in H, ordered by reverse refine-
n
S
ment. A famous result of Zaslavsky, which we will prove in Chapter 5, is that |χP (−1)| and |χP (1)| count
the number of regions and bounded regions of X, respectively.
There are additional techniques we can use for computing Möbius functions and characteristic polynomials
of lattices, particularly lattices with good structural properties (e.g., semimodular).
Definition 2.4.1. Let L be a lattice. The Möbius algebra Möb(L) is the vector space of formal C-linear
combinations of elements of L, with multiplication given by the meet operation and extended linearly. (In
particular, 1̂ is the multiplicative unit of Möb(L).)
51
The elements of L form a vector space basis of Möb(L) consisting of idempotents (elements that are their
own squares), since x ∧ x = x for all x ∈ L. For example, if L = 2[n] then Möb(L) ∼ = C[x1 , . . . , xn ]/(x21 −
2
x1 , . . . , xn − xn ), with a natural vector space basis given by squarefree monomials.
It seems as though Möb(L) could have a complicated ring structure, but actually it is quite simple.
Proposition 2.4.2. Let L be a finite lattice with n elements. For x ∈ L, define
X
εx = µ(y, x)y ∈ Möb(L).
y≤x
Then the set B = {εx : x ∈ L} is a C-vector space basis for Möb(L), with εx εy = δxy εx . In particular, Möb(L) ∼
=
Cn as rings.
Q space basis for Möb(L) as claimed. Let Cx be a copy of C with unit 1x , so that we
In particular, B is a vector
can identify C|L| with x∈L Cx . This is the direct product of rings, with multiplication 1x 1y = δxy 1x . We
claim that the C-linear map ϕ : Möb(L) → Cn given by ϕ(εx ) = 1x is a ring isomorphism. It is certainly a
vector space isomorphism, and (2.8) implies that
X X X X X
ϕ(x)ϕ(y) = ϕ εw ϕ εz = 1w 1z = 1v = ϕ(x ∧ y).
w≤x z≤y w≤x z≤y v≤x∧y
The Möbius algebra leads to useful identities that rely on translating between the “combinatorial” basis L
and the “algebraic” basis B. Some of these identities permit computation of µ(x, y) by summing Pover a clev-
erly chosen subset of [x, y], rather than the entire interval. Of course we know that µ(P ) = − x̸=1̂ µ(0̂, x)
for any poset P , but calculating µ(P ) explicitly using this formula requires a recursive computation that can
be quite inefficient. The special structure of a lattice L leads to much more streamlined expressions for µ(L).
The first of these, Weisner’s theorem (Prop. 2.4.4), reduces the number of summands substantially; it is easy
to prove and has useful consequences, but is still recursive. The second, Rota’s crosscut theorem (Thm. 2.4.9),
requires more setup but is non-recursive, which makes it a more versatile tool.
Proposition 2.4.4 (Weisner’s theorem). Let L be a finite lattice with |L| ≥ 2, and let a ∈ L \ {1̂}. Then
X
µ(x, 1̂) = 0. (2.9)
x∈L:
x∧a=0̂
52
Proof. We work in Möb(L) and calculate aε1̂ in two ways. On the one hand
X
aε1̂ = εb ε1̂ = 0.
b≤a
Now taking the coefficient of 0̂ on both sides gives (2.9), and (2.10) follows immediately.
Example 2.4.5 (The Möbius function of the partition lattice Πn ). Let a = 1|23 · · · n ∈ Πn . Then the
partitions x that show up in the sum of (2.10) are just the atoms whose non-singleton block is {1, i} for
some i > 1. For each such x, the interval [x, 1̂] ⊆ Πn is isomorphic to Πn−1 , so (2.10) gives
µ(Πn ) = − (n − 1)µ(Πn−1 )
and by induction
n
µ(Ln (q)) = (−1)n q ( 2 ) .
◀
Proof. It is sufficient to prove that (−1)r(L) µ(L) ≥ 0, since every interval in a USM lattice is USM.
Let a ∈ L \ {0̂}. Applying Weisner’s theorem to L∗ and using the fact that µ(P ) = µ(P ∗ ) (Corollary 2.2.10),
we see that X
µ(0̂, x) = 0. (2.11)
x∈L: x∨a=1̂
Now, suppose L is USM of rank n. The theorem is certainly true if n ≤ 1, so we proceed by induction.
Take a to be an atom. If x ∨ a = 1̂, then
53
so either x = 1̂, or else x is a coatom whose meet with a is 0̂. Therefore, we can solve for µ(0̂, 1̂) in (2.11) to
get X
µ(0̂, 1̂) = − µ(0̂, x).
coatoms x: x∧a=1̂
But each interval [0̂, x] is itself a USM lattice of rank n − 1, so by induction each summand has sign (−1)n−1 ,
which completes the proof.
A drawback of Weisner’s theorem is that it is still recursive; the right-hand side of (2.10) involves other
values of the Möbius function. This is not a problem for integer-indexed families of lattices {Ln } such
that every rank-k element x ∈ Ln has [0̂, x] ∼= Lk (as we have just seen), but this is too much to hope for
in general. The next result, Rota’s crosscut theorem, gives a non-recursive way of computing the Möbius
function.
Definition 2.4.8. Let L be a lattice. An upper crosscut of L is a set X ⊆ L \ {1̂} such that if y ∈ L \ X \ {1̂},
then y < x for some x ∈ X. A lower crosscut of L is a set X ⊆ L \ {0̂} such that if y ∈ L \ X \ {0̂}, then
y > x for some x ∈ X.
It would be simpler to define an upper (resp., lower) crosscut as a set that contains all coatoms (resp.,
atoms), but in practice the formulation in the previous definition is typically a convenient way to show that
a particular set is a crosscut.
Theorem 2.4.9 (Rota’s crosscut theorem). Let L be a finite lattice and X ⊂ L an upper crosscut. Then
X
µ(L) = (−1)|A| . (2.12a)
V
A⊆X: A=0̂
Proof. We prove only (2.12a); the proof of (2.12b) is dual. Fix x ∈ L and start with the following equation in
Möb(L) (recalling (2.8)): X X X
1̂ − x = εy − εy = εy .
y∈L y≤x y̸≤x
where Y = {y ∈ L : y ̸≤ x for all x ∈ X}. (Expand the sum and recall that εy εy′ = δyy′ εy .) But if X is an
upper crosscut, then Y = {1̂}, and this last equation becomes
Y X
(1̂ − x) = ε1̂ = µ(y, 1̂)y. (2.13)
x∈X y∈L
Now equating the coefficients of 0̂ on the right-hand sides of (2.13) and (2.14) yields (2.12a).
54
Corollary 2.4.10 (Möbius numbers of some lattices are boring). Let L be a lattice in which 1̂ is not a join
of atoms (for example, a distributive lattice that is not Boolean, such as almost any principal order ideal in Young’s
lattice). Then µ(L) = 0.
The crosscut theorem will be useful in studying hyperplane arrangements. Another topological application
is the following result due to J. Folkman (1966), whose proof (omitted) uses the crosscut theorem.
Theorem 2.4.11. Let L be a geometric lattice of rank r, and let P = L \ {0̂, 1̂}. Then
(
∼ Z
|µ(L)|
if i = r − 2,
H̃i (∆(P ), Z) =
0 otherwise
where H̃i denotes reduced simplicial homology. That is, ∆(P ) has the homology type of the wedge of µ(L) spheres of
dimension r − 2.
2.5 Exercises
Problem 2.1. Let P be a locally finite poset. Consider the incidence function κ ∈ I(P ) defined by
(
1 if x ⋖ y,
κ(x, y) =
0 otherwise.
Solution: 1. κn (x, y) is the number of saturated chains x, y of length exactly n, i.e., chains x = z0 ⋖ z1 ⋖ · · · ⋖
zn = y.
2. P is ranked if and only if, for all x < y, we have κn (x, y) ̸= 0 for exactly one n ∈ N.
P P
3. κ ∗ ζ(x, y) = z∈[x,y] κ(x, z)ζ(z, y) = z∈[x,y] κ(x, z) is the number of atoms in the subposet [x, y]
(equivalently, elements z such that x ⋖ z ≤ y). Likewise, ζ ∗ κ(x, y) is the number of coatoms (elements z
such that x ≤ z ⋖ y).
Problem 2.2. Prove that the Möbius function is multiplicative on direct products (i.e., µP ×Q = µP µQ in the
notation of Proposition 2.1.4) directly from the definition of µ.
Solution: Let (x, a), (y, b) ∈ P × Q with (x, a) ≤ (y, b). We wish to show that µP ×Q (x, a), (y, b) =
µP (x, y)µQ (a, b). To do this, induct on n = [(x, a), (y, b)] . The base case n = 1 occurs when x = y and
55
a = b; then the desired equation says that 1 = 1 × 1, which is true. If (x, a) < (y, b), then
X
µP ×Q (x, a), (y, b) = − µP ×Q (x, a), (z, c)
(z,c)∈P ×Q:
(x,a)≤(z,c)<(y,b)
X
=− µP (x, z)µQ (a, c) (by induction)
(z,c)∈P ×Q:
(x,a)≤(z,c)<(y,b)
X X
= − µP (x, z) µQ (a, c) − µP (x, y)µQ (a, b)
z∈P : x≤z≤y c∈Q: a≤c≤b
because at least one of the intervals [x, y], [a, b] is nontrivial, and the corresponding parenthesized sum is
zero.
Problem 2.3. Let P be a finite bounded poset and let P ∗ be its dual; recall that this means that x ≤P y if
and only if y ≤P ∗ x. Consider the vector space map F : I(P ) → I(P ∗ ) given by F (α)(y, x) = α(x, y).
(a) Show that F is an anti-isomorphism of algebras, i.e., it is a vector space isomorphism and F (α ∗ β) =
F (β) ∗ F (α).
(b) Show that F (δP ) = δP ∗ and F (ζP ) = ζP ∗ . Conclude that F (µP ) = µP ∗ and therefore that µ(P ) =
µ(P ∗ ).
Solution: (a) Define a map G : I(P ∗ ) → I(P ) in the same way. Then F ◦ G and G ◦ F are the identity
maps on I(P ∗ ) and I(P ) respectively, so both are vector space isomorphisms. To see that F is an algebra
anti-isomorphism, let α, β ∈ I(P ) and let[x, y] ∈ Int(P ), so that [y, x] ∈ Int(P ∗ ). Then
X
(F (β) ∗ F (α))(y, x) = F (β)(y, z) · F (α)(z, x)
z∈[y,x]P ∗
X
= α(x, z) · β(z, y)
z∈[x,y]P
= (α ∗ β)(x, y)
= F (α ∗ β)(y, x).
and ( (
1 x ≤P y 1 y ≤P ∗ x
F (ζP )(y, x) = ζP (x, y) = = = ζP ∗ (y, x)
0 x ̸≤P y 0 y ̸≤P ∗ x
Using these identities and the fact that F is an anti-isomorphism, we have
(and likewise if the order of convolution is reversed), so F (µP ) is the two-sided convolution inverse of ζP ∗ ,
namely µP ∗ .
56
Problem 2.4. (Based on an observation by Mark Denker) Prove that
X
pG (t) = ck (G)t(t − 1)(t − 2) · · · (t − k + 1)
k
where ck (G) is the number of ways of properly coloring G using exactly k colors (i.e., proper colorings that
are surjective functions κ : V (G) → [k]).
Problem 2.5. A set partition in Πn is a noncrossing partition (NCP) if its associated equivalence relation ∼
satisfies the following condition: for all i < j < k < ℓ, if i ∼ k and j ∼ ℓ then i ∼ j ∼ k ∼ ℓ. The set of all
NCPs of order n is denoted NCn . Ordering by reverse refinement makes NCn into a subposet of the partition
lattice Πn . Note that NCn = Πn for n ≤ 3 (the smallest partition that is not noncrossing is 13|24 ∈ Π4 ). NCPs
can be represented pictorially by chord diagrams. The chord diagram of ξ = 1|2 5|3|4|6 8 12|7|9|10 11 ∈ NC12
is shown in Figure 2.1(a).
12 11’ 12 12’
(a) 11 1 (b) 11 1
10’ 1’
10 2 10 2
9’ 2’
9 3 9 3
8’ 3’
8 4 8 4
7’ 4’
7 5 7 5
6 6’ 6 5’
1 2n
therefore, ncn is the nth Catalan number Cn = n+1 n . (You can formally define nc0 = 1, and
establish the recurrence for n ≥ 1.)
(c) Prove that the operation of Kreweras complementation is an anti-automorphism of NCn . To define the
Kreweras complement K(π) of π ∈ NCn , start with the chord diagram of π and insert a point labeled i′
between the points i and i + 1 (mod n) for i = 1, 2, . . . , n. Then a, b lie in the same block of K(π) if it is
possible to walk from a′ to b′ without crossing an arc of π. For instance, the Kreweras complement of
the noncrossing partition ξ ∈ NC12 shown above is K(ξ) = 1 5 12|2 3 4|6 7|8 9 11|10 (see Figure 2.1(b)).
(d) Use Weisner’s theorem to prove that µ(NCn ) = (−1)n−1 Cn−1 for all n ≥ 1.
The characteristic polynomial of NCn satisfies a version of the Catalan recurrence. For details see [LS00]
(this might make a good end-of-semester project).
Solution: (a) NCn is certainly a bounded subposet of Πn with the same top and bottom elements, and every
relation in NCn is a relation in Πn . Recall that for π, σ ∈ Πn , the blocks of τ = π ∧ σ are the nonempty
intersections of blocks of π with blocks of σ (see Example 1.2.2). In particular, every two blocks of τ come
from either different blocks of π or different blocks of σ, so if π and σ are noncrossing then so is τ . It follows
that NCn is a bounded meet-semilattice, hence a lattice. It is not a sublattice of Πn because the join (in Πn )
57
of two noncrossing partitions need not be noncrossing. For example, if π = 13|2|4 and σ = 24|1|3, then their
join in Πn is 13|24, which is not noncrossing; their join in NCn is 1234. (By the way, the join of two NCPs
can be obtained by repeatedly merging any blocks that either intersect or cross — this is easiest to calculate
visually by superimposing the chord diagrams.)
We will show that NCn inherits the rank function of Πn , namely r(π) = n − |π|. To do this, we need to show
that every covering relation in NCn is a covering relation in Πn . That is, if σ ⋗ π then π is obtained from σ
by splitting one block into two pieces.
First, I claim that every π ∈ NCn has some block that is an interval. Let k = |π|. If k = 1 then the conclusion
is trivial; otherwise, label the blocks of π so that min(π1 ) < · · · < min(πk ). If [min(πk ), max(πk )] contains
any element from another block πi , then πi and πk cross. Therefore, πk is an interval. A consequence is that
π lies below the coatom with two blocks πk and [n] \ πk .
By extension, if π ≤ σ and σi is a block of σ that is partitioned into two or more blocks of π, then there is
some such block πk such that πk = [min(πk ), max(πk )] ∩ σi . Therefore, splitting σi as πk ∪· (σi \ πk ) produces
a partition ρ such that σ ⋗ ρ ≥ π. This completes the proof.
(b) If n ≤ 3 then nc1 = 1, nc2 = 2, nc3 = 5, which are the first three Catalan numbers. For larger n, consider
ξ ∈ NCn . It is possible that n is a singleton block in ξ, in which case there are ncn−1 possibilities for ξ.
Otherwise, let k ∈ [n − 1] be the smallest number in the block of ξ containing n. Contract the edge kn: then
ξ can be specified by a pair (ξ ′ , ξ ′′ ) ∈ NC(k − 1) × NC(n − k), where ξ ′′ is interpreted as a NCP on [k + 1, n].
(An example is shown below, with n = 12 and k = 6.) The Catalan recurrence follows.
12
11 1 10 11 1 2
10 2
9 3 9 12 3
8 4
7 5 8 7 5 4
6
ξ ξ 00 ξ0
(c) First, note that K(K(π)) is the NCP obtained by subtracting 1 (mod n) from every element of [n] — see
example below. This is a bijection, so K is as well (although it is not an involution).
58
11’ 12 12’ 11 11’ 12
11 1 10’ 12’
10’ 1’ 10 1
10 2 9’ 1’
9’ 2’ ξ 9 2
9 3 K(ξ) 8’ 2’
8’ 3’ K(K(ξ)) 8 3
8 4 7’ 3’
7’ 4’ 7 4
7 5 6’ 4’
6’ 6 5’ 6 5’ 5
Second, if π is a coarsening of σ, then every arc of σ is an arc of π, which means that two numbers in the
same block of K(π) are in the same block of K(σ), so K(σ) is a coarsening of K(π). Therefore Kreweras
complementation reverses order in NCn .
(d) Let mn = µ(NCn ). We proceed by induction on n. For n ≤ 3 we can compute directly m1 = 1, m2 = −1,
m3 = 2.
For larger n, let a = {{1, 2, . . . , n − 1}, {n}} ∈ NCn . Note that x ∧ a = 0̂ if and only if all of 1, . . . , n − 1 are
singleton blocks in x. For 1 ≤ k ≤ n − 1, let bk be the partition with {k, n} as the only non-singleton block.
Then [bk , 1̂] ∼
= NCk × NCn−k for the same reason as in part (b). Therefore
n
X
mn = µ(NCn ) = − µ(bk , 1̂) (by Weisner’s theorem)
k=1
n−1
X
=− mk mn−k (since µ is multiplicative on direct products)
k=1
n−1
X
=− (−1)k−1 Ck−1 (−1)n−k−1 Cn−k−1 (by induction)
k=1
n−2
X
= −(−1)n Cj Cn−j−2 (reindexing j = k − 1)
j=0
n−2
X
= (−1)n−1 Cn−2 + Cj Cn−j−2
j=1
n−1
= (−1) Cn−1 (by the Catalan recurrence).
Problem 2.6. This problem is about how far Proposition 2.4.2 can be extended. Suppose that R is a commu-
tative C-algebra of finite dimension n as a C-vector space, and that x1 , . . . , xn ∈ R are linearly independent
idempotents (i.e., x2i = xi for all i). Prove that R ∼
= Cn as rings.
Solution: This solution is due to Darij Grinberg. First note that each 1 − xi is an idempotent as well, since
(1 − xi )2 = 1 − 2xi + x2i = 1 − 2xi + xi = 1 − xi . Also, xi (1 − xi ) = xi − x2i = 0. Moreover, the product of
59
two idempotents is an idempotent. Therefore, for each S ⊆ [n], we can define an idempotent
Y Y
eS = ( xi )( (1 − xi )).
i∈S i∈[n]\S
These are mutually orthogonal, since if S ̸= T , say j ∈ S△T , then the factor xi (1 − xi ) appears in xS xT . In
addition, they span R, because every xi is a linear combination of them, specifically
X Y
eS = xi (xj + (1 − xj )) = xi .
S⊆[n]: i∈S j̸=i
P
In addition, the set of nonzero elements eS is linearly independent, for if S λS eS = 0 for some scalars λS ,
then multiplying both sides by eT for some given T ⊆ [n] yields λT eT = 0 (in view of the orthogonality
of the idempotents eS ) and thus either λT = 0 or eT = 0 (since we are working over a field). But, of
course, these nonzero elements span R (since all the eS span R), so they form a basis. Thus, R has a basis
consisting of orthogonal idempotents. This basis must have n elements (since dim R = n), and thus R ∼ = Cn
as C-algebras.
Problem 2.7. The q-binomial coefficient is the rational function
(q n − 1)(q n − q) · · · (q n − q k−1 )
n
= .
k q (q k − 1)(q k − q) · · · (q k − q k−1 )
(a) Check that setting q = 1 (after canceling out common terms), or equivalently applying limq→1 , recov-
ers the ordinary binomial coefficient nk .
(b) Prove the q-Pascal identities: for n ≥ 1 and all k, we have
n−1 n−1 n−1 n−1
n n
= qk + and = + q n−k .
k q k q k−1 q k q k q k−1 q
(c) Prove that nk q is the generating function for partitions that fit into the rectangle Rk,n−k with n − k
(This result has a surprising application to the topology of algebraic varieties; see §11.4.)
(d) (Stanley, EC1, 2nd ed., 3.119) Prove the q-binomial theorem:
n−1 n
n k
(−1)k q (2) xn−k .
Y X
(x − q k ) =
k q
k=0 k=0
Solution: (a) For n ∈ N>0 , write [n]q = (1 − q n )/(1 − q) = q n−1 + q n−2 + · · · + q + 1. (This polynomial is
called the q-analogue of n, pronounced “n base q.”) Thus
k−1
Y (q n − q j ) k−1
Y (q n−j − 1) k−1
Y (q − 1)[n − j]q k−1
Y [n − j]q
n
= = = = .
k q j=0 (q k − q j ) j=0
(q k−j − 1) j=0
(q − 1)[k − j]q j=0
[k − j]q
60
Now setting q = 1 gives
n(n − 1) · · · (n − k − 1)
n! n
= = .
k(k − 1) · · · 1 (n − k)!k! k
(b) The two proofs are similar; here’s the first one.
Y [n − 1 − j]q k−2
k−1
n−1 n−1 Y [n − 1 − j]q
k
q + = qk +
k q k−1 q j=0
[k − j]q j=0
[k − 1 − j]q
k−1 k−1
Y [n − 1 − j]q Y [n − j]q
= qk +
j=0
[k − j]q j=1
[k − j]q
k−1
Y [n − j]q [n − k]q
k
= q +1
j=1
[k − j]q [k]q
k−1
Y [n − j]q q k (q n−k − 1) + (q k − 1)
=
j=1
[k − j]q qk − 1
k−1
Y [n − j]q q n − 1
=
j=1
[k − j]q qk − 1
k−1
Y [n − j]q [n]q k−1
Y [n − j]q
n
= = = .
j=1
[k − j]q [k] q j=0
[k − j] q k q
(c) When n = k or k = 0, the result is trivially true: it asserts that 1 = 1. Thus, suppose that 0 < k < n, and
assume inductively that the result holds for all smaller values of n and all values of k.
Consider a partition λ ≤ R(k, n − k). The first possibility is that λ1 = k, in which case removing the first
row produces a partition µ ⊆ R(k, n − k − 1) with |µ| = |λ| − k, and the correspondence between λ and µ
is a bijection. On the other hand, if λ1 < k, then λ fits into the rectangle R(k − 1, n − k), and again this is a
bijection. Therefore,
X X X
q |λ| = q k q |λ| + q |λ|
λ⊆R(k,n−k) µ⊆R(k,n−k−1) λ⊆R(k−1,n−k)
n−1 n−1
= qk + (by induction)
k q k−1 q
n
= (by the q-Pascal recurrence).
k q
(d) Let e1 , . . . , en be a basis for V . Then a linear transformation ϕ : V → X is specified by its values on
ei . For ϕ to be one-to-one, it is necessary and sufficient to choose ϕ(e1 ), . . . , ϕ(en ) in order so that ϕ(ei+1 ) ̸∈
Fq ⟨ϕ(e1 ), . . . , ϕ(ei )⟩; the number of choices is therefore (x − 1)(x − q)(x − q 2 ) · · · (x − q n−1 ), which is the
left-hand side of the desired equation. This is the characteristic polynomial of the subspace lattice Ln (q)
(see Example 2.4.6). By definition of the characteristic polynomial,
n−1 n
k
#{Z ⊆ V : dim Z = k}(−1)k q (2) xn−k .
Y X X
k n−dim Z
(x − q ) = µ(0, Z)x = (2.15)
k=0 Z⊆V k=0
61
There are (q n − 1)(q n − q) · · · (q n − q k−1 ) ordered lists of k linearly independent vectors in Fnq , and each one
k k k
spans a k-dimensional subspace. On the other hand, each such subspace nhas (q − 1)(q − q) · · · (q − 1)
ordered bases. It follows that the number of k-dimensional subspaces is k q , so that (2.15) becomes
n−1 n
n k
(−1)k q (2) xn−k
Y X
k
(x − q ) =
k q
k=0 k=0
which is the q-binomial theorem. (Setting q = 1 recovers the ordinary binomial theorem.) Note also that
since there is a bijection between k- and (n − k)-dimensional subspaces of Fnq (e.g., orthogonal complement
with respect to any non-degenerate bilinear form, such as the standard one), it follows that
n n
=
k q n−k q
(which can also be checked from the definition as rational functions, though the derivation is mildly te-
dious).
Problem 2.8. (Stanley, EC1, 3.129) Here is a cute application of combinatorics to elementary number theory.
Let P be a finite poset, and let P̂ = P ∪ {0̂, 1̂}. Suppose that P has a fixed-point-free automorphism
σ : P → P of prime order p; that is, σ(x) ̸= x and σ p (x) = x for all x ∈ P . Prove that µP̂ (0̂, 1̂) ≡ −1
(mod p). What does this say in the case that P̂ = Πp ?
Solution: Since σ is an automorphism, the Möbius function is constant on its orbits. Therefore, for an orbit
O, we may write µ(O) for the value of µ(0̂, x) for any x ∈ O. Every orbit has cardinality p, so
X X
µP̂ (0̂, 1̂) = − µ(0̂, x) = − 1 − µ(0̂, x)
x∈P ∪{0̂} x∈P
X X
= −1− µ(0̂, x)
orbits O x∈O
X
= −1−p µ(O) ≡ −1 (mod p)
orbits O
as desired. (Alternate proof: by Philip Hall’s Theorem, µP (0̂, 1̂) = k≥0 (−1)k ck , where ck is the number
P
of strictly increasing chains 0̂ = x0 < x1 < · · · < xk = 1̂. Now c0 = 0 (since P has at least two elements)
and c1 = 1. Meanwhile, σ takes k-chains to k-chains, and for every k ≥ 2 each orbit has size p, so ck ≡ 0
(mod p) and the desired congruence follows.)
a fact known as Wilson’s theorem. (You can actually get rid of the (−1)p−1 factor; either p = 2 when −1 ≡ 1
(mod p), or else p is odd and (−1)p−1 = 1.) Note also that primality is necessary as well as sufficient — if n
is composite then (n − 1)! is divisible by n.
62
Chapter 3
Matroids
The motivating example of a geometric lattice is the lattice of flats of a finite set E of vectors. The underlying
combinatorial data of this lattice can be expressed in terms of the rank function, which says the dimension
of the space spanned by every subset of E. However, there are many other equivalent ways to describe
the “combinatorial linear algebra” of a set of vectors: the family of linearly independent sets; the family
of sets that form bases; which vectors lie in the span of which sets; etc. Each of these data sets defines the
structure of a matroid on E. Matroids can also be regarded as generalizations of graphs, and are important
in combinatorial optimization as well. A standard reference on matroid theory is [Oxl92], although I first
learned the basics of the subject from an unusual (but very good) source, namely chapter 3 of [GSS93].
Conventions: Unless otherwise specified, E always denotes a finite set. We will be doing a lot of adding
elements to and removing elements e from sets A, so for convenience we define A + e = A ∪ {e} and
A − e = A \ {e}.
Definition 3.1.1. Let E be a finite set. A closure operator on E is a map 2E → 2E , written A 7→ Ā, such that
(i) A ⊆ Ā;
(ii)  = ; and
(iii) if A ⊆ B, then Ā ⊆ B̄
63
Definition 3.1.3. A matroid closure operator is a closure operator that satisfies the exchange property:
A matroid M is a set E (the “ground set”) together with a matroid closure operator on E. A matroid is
simple if the empty set and all singleton sets are closed.
Example 3.1.4. Vector matroids. Let V be a vector space over a field k, and let E ⊆ V be a finite set. Then
A 7→ Ā := kA ∩ E
is a matroid closure operator on E. It is easy to check the conditions for a closure operator. To check
condition (3.2), if e ∈ A + e′ , then there is a linear equation
X
e = ce′ e ′ + ca a
a∈A
where ce′ and all the ca are scalars in k. The condition e ̸∈ Ā implies that ce′ ̸= 0 in any equation of this
form. Therefore, the equation can be rewritten to express e′ as a linear combination of the vectors in A + e,
obtaining (3.2). A matroid arising in this way (or, more generally, isomorphic to such a matroid) is called a
vector matroid, linear matroid or representable matroid.1 ◀
A vector matroid records information about linear dependence (i.e., which vectors belong to the linear
spans of other sets of vectors) without having to worry about the actual coordinates of the vectors. More
generally, a matroid can be thought of as a combinatorial, coordinate-free abstraction of linear dependence
and independence. Note that a vector matroid is simple if none of the vectors is zero (so that ¯∅ = ∅) and if
no vector is a scalar multiple of another (so that all singleton sets are closed).
The following theorem says that simple matroids and geometric lattices are essentially the same things. In
Rota’s language, they are “cryptomorphic”: their definitions look very different, but they carry the same in-
formation. We will see many more ways to axiomatize the same information: rank functions, independence
systems, basis systems, etc. Working with matroids requires a solid level of comfort with the cryptomor-
phisms between the various definitions of a matroid.
Theorem 3.2.1. 1. Let M be a simple matroid with finite ground set E, and L(M ) its lattice of flats. Then L(M )
is a geometric lattice, under the operations F ∧ G = F ∩ G, F ∨ G = F ∪ G. W
2. Let L be a geometric lattice and let E be its set of atoms. Then the function A 7→ A = {e ∈ E : e ≤ A} is a
matroid closure operator for a simple matroid on E.
3. These constructions are mutual inverses.
Proof. (1) Recall that L(M ) is a lattice by Proposition 3.1.2. By definition of a simple matroid, the bottom
element is ∅ and the atoms are the singleton subsets of E. Every flat is the join of the atoms corresponding
to its elements, so the lattice L(M ) is atomic.
64
Proof. Suppose that there is a flat G such that
F ⊊ G ⊆ F ∨ e = F + e. (3.3)
Let e′ ∈ G \ F . Then e′ ∈ F + e, so the exchange axiom (3.2) implies e ∈ F + e′ , which in turn implies that
F ∨ e ⊆ F ∨ e′ ⊆ G. Hence the ⊆ in (3.3) is actually an equality. We have shown that there are no flats
strictly between F and F ∨ e, proving the claim.
Of course, if F ⋖ G then G = F ∨ e for any e ∈ G \ F . So the covering relations in L(M ) are precisely of this
form.
Suppose now that F and G are incomparable and that G ⋗ F ∧ G. Then G is of the form (F ∧ G) ∨ e, and
we can take e to be any element of G \ F . In particular F < F ∨ e, so by Lemma 3.2.2, F ⋖ F ∨ e. Moreover,
F ∨ G = F ∨ ((F ∧ G) ∨ e) = (F ∨ e) ∨ (F ∧ G) = F ∨ e.
We have just proved that L(M ) is semimodular. Here is the diamond picture (cf. (1.7)):
F ∨G=F ∨e
•
F G = (F ∧ G) ∨ e
•
F ∧G
Recall that if L is semimodular and x, e ∈ L with e an atom and x ̸≥ e (so that x < x ∨ e), then in fact
x ⋖ x ∨ e, because
r(x ∨ e) − r(x) ≤ r(e) − r(x ∧ e) = 1 − 0 = 1.
Accordingly,
W let A ⊆ E and let e, f ∈ E \ A. Suppose that e ∈ A + f ; we must show that f ∈ A + e. Let
x = A ∈ L. Then
65
Example 3.2.3. Let E = {a, b, c, d, e}. The subfamily of 2E given by
regarded as a poset under inclusion, is a geometric lattice of rank 3 (feel free to check this yourself). In the
displayed equation, the elements of L are grouped by rank. The associated matroid closure operator has,
for example, ad = ad, ac = abc = abc, ace = abcde. ◀
In view of Theorem 3.2.1, we can describe a matroid on ground set E by the function A 7→ r(Ā), where r
is the rank function of the associated geometric lattice. It is standard to abuse notation by calling this
function r as well. Formally:
Definition 3.2.4. A matroid rank function on E is a function r : 2E → N satisfying the following conditions
for all A, B ⊆ E:
If r is a matroid rank function on E, then the corresponding matroid closure operator is given by
A = {e ∈ E : r(A + e) = r(A)}.
Moreover, this matroid is simple if and only if r(A) = |A| whenever |A| ≤ 2.
Conversely, if A 7→ Ā is a matroid closure operator on E, then the corresponding matroid rank function r is
It is easy to check that this satisfies the conditions of Definition 3.2.4. The corresponding matroid is called
the uniform matroid Uk (n). Its closure operator is
(
A if |A| < k,
A=
E if |A| ≥ k.
So the flats of M are the sets of cardinality < k, as well as E itself. Therefore, the lattice of flats looks like
a Boolean lattice 2[n] that has been truncated at the kth rank: that is, all elements of rank ≥ k have been
deleted and replaced with a single 1̂. For n = 3 and k = 2, this lattice is M5 . For n = 4 and k = 3, the Hasse
diagram is as shown below.
1234
12 13 14 23 24 34
1 2 3 4
66
If E is a set of n vectors in general position in kk , then the corresponding linear matroid is isomorphic to
Uk (n). This sentence is tautological, in the sense that it can be taken as a definition of “general position”.
If k is infinite and the points are chosen randomly (in some reasonable measure-theoretic sense), then L(E)
will be isomorphic to Uk (n) with probability 1. On the other hand, k must be sufficiently large (in terms
of n) in order for kk to have n points in general position: for instance, U2 (4) cannot be represented as a
matroid over F2 simply because F22 contains only three nonzero vectors. ◀
In general, every definition of “matroid” (and there are several more coming) will induce a corresponding
equivalent for isomorphic.
Let G be a finite graph with vertices V and edges E. For convenience, we will write e = xy to mean “e is an
edge with endpoints x, y”. This notation does not exclude the possibility that e is a loop (i.e., x = y) or that
some other edge might have the same pair of endpoints.
Definition 3.3.1. For each subset A ⊆ E, the corresponding induced subgraph of G is the graph G|A with
vertices V and edges A. The graphic matroid or complete connectivity matroid M (G) on E is defined by
the closure operator
Equivalently, an edge e = xy belongs to Ā if there is a path between x and y consisting of edges in A (for
short, an A-path). For example, in the graph, 14 ∈ Ā because {12, 24} ⊆ A.
1 1 1
2 3 2 3 2 3
4 5 4 5 4 5
G = (V, E) A Ā
Proof. It is easy to check that A ⊆ Ā for all A, and that A ⊆ B =⇒ Ā ⊆ B̄. If e = xy ∈ , then x, y can
be joined by an Ā-path P , and each edge in P can be replaced with an A-path, giving an A-path between x
and y.
67
Finally, suppose e = xy ̸∈ Ā but e ∈ A + f . Let P be an (A + f )-path from x to y. Then f ∈ P (because there
is no A-path from x to y) and P + e is a cycle. Deleting f produces an (A + e)-path between the endpoints
of f . (See Figure 3.1.)
e P f
Figure 3.1: The closure axiom for a graphic matroid. Here A consists of all edges shown except e and f ;
neither e nor f belongs to Ā, but e ∈ A + f and f ∈ A + e.
1. r(B) = r(A);
2. B is acyclic;
3. |B| = |V | − c, where c is the number of connected components of A.
The flats of M (G) correspond to the subgraphs of G in which every component is an induced subgraph
of G. In other words, the geometric lattice corresponding to the graphic matroid M (G) is precisely the
connectivity lattice K(G) introduced in Example 1.2.3.
Example 3.3.4. If G is a forest (a graph with no cycles), then no two vertices are joined by more than one
path. Therefore, every edge set is a flat, and M (G) ∼
= Un (n). ◀
Example 3.3.5. If G is a cycle of length n, then every edge set of size < n − 1 is a flat, but the closure of a set
of size n − 1 is the entire edge set. Therefore, M (G) ∼
= Un−1 (n). ◀
Example 3.3.6. If G = Kn (the complete graph on n vertices), then a flat of M (G) is the same thing as an
equivalence relation on [n]. Therefore, M (Kn ) is naturally isomorphic to the partition lattice Πn . ◀
In addition to rank functions, lattices of flats, and closure operators, there are many other equivalent ways
to define a matroid on a finite ground set E. In the fundamental example of a linear matroid M , some of
these definitions correspond to linear-algebraic notions such as linear independence and bases.
2 This terminology can cause confusion. By definition a subgraph H of G is spanning if V (H) = V (G), but not every acyclic
spanning subgraph is a spanning forest. A more accurate term would be “maximal forest”.
68
Definition 3.4.1. A (matroid) independence system I is a family of subsets of E such that
(I1) ∅ ∈ I ;
(I2) if I ∈ I and I ′ ⊆ I, then I ′ ∈ I ;
(I3) (“Donation”) if I, J ∈ I and |I| < |J|, then there is some x ∈ J \ I such that I ∪ x ∈ I .
Note that conditions (I1) and (I2) say that I is an abstract simplicial complex on E (see Example 1.1.11).
If E is a finite subset of a vector space, then the linearly independent subsets of E form a matroid indepen-
dence system. Conditions (I1) and (I2) are clear. For (I3), the span of J has greater dimension than that of
I, so there must be some x ∈ J outside the span of I, and then I ∪ x is linearly independent.
The next lemma generalizes the statement that any linearly independent set of vectors can be extended to
a basis of any space containing it.
Lemma 3.4.2. Let I be a matroid independence system on E. Suppose that I ∈ I and I ⊆ X ⊆ E. Then I can be
extended to a maximum independent subset of X.
Proof. If I already has maximum cardinality then we are done. Otherwise, let J be a maximum independent
subset of X. Then |J| > |I|, so by (I3) there is some x ∈ J \ I with I ∪ x independent. Replace I with I ∪ x
and repeat.
The argument shows also that for every X ⊆ E, all maximal independent subsets (or bases) of X have
the same cardinality (so there is no irksome difference between “maximal” and “maximum”). In simplicial
complex terms, every induced subcomplex of I is pure — an induced subcomplex is something of the form
I |X = {I ∈ I : I ⊆ X}, for X ⊆ E, and “pure” means that all maximal faces have the same cardinality.
This condition actually characterizes matroid independence complexes; we will take this up again in §6.5.
A matroid independence system records the same combinatorial structure on E as a matroid rank function:
is an independence system.
2. If I is an independence system on E, then
Proof. Part 1: Let r be a matroid rank function on E and define I as in (3.5a). First, r(I) ≤ |I| for all I ⊆ E,
so (I1) follows immediately. Second, suppose I ∈ I and I ′ ⊆ I; say I ′ = {x1 , . . . , xk } and I = {x1 , . . . , xn }.
Consider the “flag” (nested family of subsets)
∅ ⊊ {x1 } ⊊ {x1 , x2 } ⊊ · · · ⊊ I ′ ⊊ · · · ⊊ I.
The rank starts at 0 and increases at most 1 each time by submodularity. But since r(I) = |I|, it must
increase by exactly 1 each time. In particular r(I ′ ) = k = |I ′ | and so I ′ ∈ I , establishing (I2).
69
To show (I3), let I, J ∈ I with |I| < |J| and let J \ I = {x1 , . . . , xn }. If n = 1 then J = I + x1 and there is
nothing to show. Now suppose that n ≥ 2 and r(I + xk ) = r(I) for every k ∈ [n]. By submodularity,
and equality must hold throughout. But then r(I ∪ J) = r(I) < r(J), which is a contradiction.
Part 2: Now suppose that I is an independence system on E, and define a function r : 2E → Z as in (3.5b).
It is immediate from the definition that r(A) ≤ |A| and that A ⊆ B implies r(A) ≤ r(B) for all A, B ∈ I .
To prove submodularity, let A, B ⊆ E and let I be a basis of A ∩ B. By Lemma 3.4.2, we can extend I to
a basis J of A ∪ B. Note that no element of J \ I can belong to both A and B, otherwise I would not be a
maximal independent set in A ∩ B. So we have the following Venn diagram:
A B
I
J
Moreover, J ∩ A and J ∩ B are independent subsets of A and B respectively, but not necessarily maximal,
so
r(A ∪ B) + r(A ∩ B) = |I| + |J| = |J ∩ A| + |J ∩ B| ≤ r(A) + r(B).
If M = M (G) is a graphic matroid, the associated independence system I is the family of acyclic edge sets
in G. To see this, notice that if A is a set of edges and e ∈ A, then r(A − e) < r(A) if and only if deleting e
breaks a component of G|A into two smaller components (so that in fact r(A − e) = r(A) − 1). This is
equivalent to the condition that e belongs to no cycle in A. Therefore, if A is acyclic, then deleting its edges
one by one gets you down to ∅ and decrements the rank each time, so r(A) = |A|. On the other hand, if A
contains a cycle, then deleting any of its edges won’t change the rank, so r(A) < |A|.
Here’s what the “donation” condition (I3) means in the graphic setting. Suppose that |V | = n, and let c(H)
denote the number of components of a graph H. If I, J are acyclic edge sets with |I| < |J|, then
and there must be some edge e ∈ J whose endpoints belong to different components of G|I ; that is, I + e is
acyclic.
The bases of M (the maximal independent sets) provide another way of defining a matroid.
Definition 3.4.4. A (matroid) basis system on E is a nonempty family B ⊆ 2E such that for all B, B ′ ∈ B,
(B1) |B| = |B ′ |;
(B2) For all e ∈ B \ B ′ , there exists e′ ∈ B ′ \ B such that (B − e) + e′ ∈ B;
70
(B2′ ) For all e ∈ B \ B ′ , there exists e′ ∈ B ′ \ B such that (B ′ + e) − e′ ∈ B.
In fact, given (B1), the conditions (B2) and (B2′ ) are equivalent, although this require some proof (Prob-
lem 3.2).
For example, if S is a finite set of vectors spanning a vector space V , then the subsets of S that are bases for
V all have the same cardinality (namely dim V ) and satisfy the basis exchange condition (B2).
If G is a graph, then the bases of M (G) are its spanning forests, i.e., its maximal acyclic edge sets. If G is
connected (which, as we will see, we may as well assume when studying graphic matroids) then the bases
of M (G) are its spanning trees.
B B0
Here is the graph-theoretic interpretation of (B2). Let G be a connected graph, let B, B ′ be spanning trees,
and let e ∈ B \ B ′ . Then B − e has exactly two connected components. Since B ′ is connected, it must have
some edge e′ with one endpoint in each of those components, and then B − e + e′ is a spanning tree. See
Figure 3.2.
B B\e B0
Figure 3.2: An example of basis axiom (B2) in a graphic matroid. The green edges are the possibilities for e′
such that B\e + e′ is a spanning tree.
As for (B2′ ), if e ∈ B \ B ′ , then B ′ + e must contain a unique cycle C (formed by e together with the unique
path P in B ′ between the endpoints of e). Deleting any edge e′ ∈ P will produce a spanning tree, and there
must be at least one such edge e′ ̸∈ B (otherwise B contains the cycle C). See Figure 3.3.
If G is a graph with edge set E and M = M (G) is its graphic matroid, then
I = {A ⊆ E : A is acyclic},
B = {A ⊆ E : A is a spanning forest of G}.
71
e
∗ ∗
B B0 B0 ∪ e
Figure 3.3: An example of basis axiom (B2′ ) in a graphic matroid. The path P is shown in green. The edges
of P \ B, marked with stars, are valid choices for e′ .
I = {A ⊆ S : A is linearly independent},
B = {A ⊆ S : A is a basis for span(S)}.
1. If I is an independence system onSE, then the family of maximal elements of I is a basis system.
2. If B is a basis system, then I = B∈B 2B is an independence system.
3. These constructions are mutual inverses.
The proof is left as an exercise. We already have seen that an independence system on E is equivalent to a
matroid rank function; Proposition 3.4.5 asserts that a basis system provides the same structure on E. Bases
turn out to be especially convenient for describing fundamental operations on matroids such as duality,
direct sum, and deletion/contraction (all of which are coming soon).
Instead of specifying the bases (maximal independent sets), a matroid can be defined by its minimal depen-
dent sets, which are called circuits. These too can be axiomatized:
Definition 3.4.6. A (matroid) circuit system on E is a family C ⊆ 2E such that, for all C, C ′ ∈ C ,
(C1) ∅ ̸∈ C :
(C2) C ⊈ C ′;
(C3) For all e ∈ C ∩ C ′ , the set (C ∪ C ′ ) − e contains an element of C .
In a linear matroid, the circuits are the minimal dependent sets of vectors. Indeed, if C, C ′ are such sets and
e ∈ C ∩ C ′ , then we can find two expressions for e as nontrivial linear combinations of vectors in C and in
C ′ , and equating these expressions and eliminating e shows that (C ∪ C ′ ) − e is dependent, hence contains
a circuit.
In a graph, if two cycles C, C ′ meet in a (non-loop) edge e = xy, then C − e and C ′ − e are paths between x
and y, so concatenating them forms a closed path. This path is not necessarily itself a cycle, but must contain
some cycle.
Proposition 3.4.7. Let E be a finite set.
72
In other words, the circuits are the minimal nonfaces of the independence complex (hence they correspond
to the generators of the Stanley-Reisner ideal; see Defn. 6.3.1). The proof is left as an exercise.
The final definition of a matroid is different from what has come before, and gives a taste of the importance
of matroids in combinatorial optimization.
Let E be a finite set and let ∆ be an abstract simplicial complex on E (see Definition 3.4.1). Let w : E → R≥0
P a function, which we regard as assigning weights to the elements of E, and for A ⊆ E, define w(A) =
be
e∈A w(e). Consider the problem of maximizing w(A) over all subsets A ∈ ∆; the maximum will certainly
be achieved on a facet. A naive approach to find a maximal-weight A, which may or may not work for a
given ∆ and w, is the following “greedy” algorithm (known as Kruskal’s algorithm):
1. Let A = ∅.
2. If A is a facet of ∆, stop.
Otherwise, find e ∈ E \ A of maximal weight such that A + e ∈ ∆ (if there are several such e, pick one
at random), and replace A with A + e.
3. Repeat step 2 until A is a facet of ∆.
Proposition 3.4.8. ∆ is a matroid independence system if and only if Kruskal’s algorithm produces a facet of maximal
weight for every weight function w.
The proof is left as an exercise, as is the construction of a simplicial complex and a weight function for
which the greedy algorithm does not produce a facet of maximal weight. This interpretation can be useful
in algebraic combinatorics; see Example 9.19.2 below.
The motivating example of a matroid is a finite collection of vectors in Rn . What if we work over a different
field? What if we turn this question on its head by specifying a matroid M purely combinatorially and then
asking which fields give rise to vector sets whose matroid is M ?
Definition 3.5.1. Let M be a matroid and V a vector space over a field k. A set of vectors S ⊆ V represents
or realizes M over k if the linear matroid M (S) associated with S is isomorphic to M .
73
For example:
• The matroid U2 (3) is representable over any field F. Set S = {(1, 0), (0, 1), (1, 1)}; any two of these
vectors form a basis of F2 .
• If k has at least three elements, then U2 (4) is representable, by, e.g., S = {(1, 0), (0, 1), (1, 1), (1, a)}.
where a ∈ k \ {0, 1}. Again, any two of these vectors form a basis of k2 .
• On the other hand, U2 (4) is not representable over F2 , because F22 doesn’t contain four nonzero ele-
ments.
More generally, suppose that M is a simple matroid with n elements (i.e., the ground set E has |E| = n) and
rank r (i.e., every basis of M has size r) that is representable over the finite field Fq of order q. Then each
element of E must be represented by some nonzero vector in Frq , and no two vectors can be scalar multiples
of each other. Therefore,
qr − 1
n≤ .
q−1
Example 3.5.2. The Fano plane. Consider the affine point configuration with 7 points and 7 lines (one of
which looks like a circle), as shown:
6 4
7
1 3
2
This point configuration cannot be represented over R. If you try to draw seven non-collinear points in
R2 such that the six triples 123, 345, 156, 147, 257, 367 are each collinear, then 246 will not be collinear —
try it. The same thing will happen over any field of characteristic ̸= 2. On the other hand, over a field
of characteristic 2, if the first six triples are collinear then 246 must be collinear. The configuration can be
explicitly represented over F2 by the columns of the matrix
1 1 0 0 0 1 1
0 1 1 1 0 0 1 ∈ (F2 )3×7
0 0 0 1 1 1 1
for which each of the seven triples of columns listed above is linearly dependent, and that each other
triple is a column basis. (Note that over R, the submatrix consisting of columns 2,4,6 has determinant 2.)
The resulting matroid is called the Fano plane or Fano matroid. Note that each line in the Fano matroid
corresponds to a 2-dimensional subspace of F32 .
Viewed as a matroid, the Fano plane has rank 3. Its bases are the 73 − 7 = 28 noncollinear triples of points.
Its circuits are the seven collinear triples and their complements (known as ovals). For instance, 4567 is an
oval: it is too big to be independent, but on the other hand every three-element subset of it forms a basis (in
particular, is independent), so it is a circuit.
74
The Fano plane is self-dual in the sense of discrete geometry3 : the lines can be labeled 1, . . . , 7 so that point i
lies on line j if and only if point j lies on line i. Here’s how: recall that the points and lines of the Fano plane
correspond respectively to 1- and 2-dimensional subspaces of F32 , and assign the same label to orthogonally
complementary spaces under the standard inner product. ◀
Example 3.5.3 (Finite projective planes). Let q ≥ 2 be a positive integer. A projective plane of order q
consists of a collection P of points and a collection L of lines, each of which is a subset of P , such that:
• |P | = |L| = q 2 + q + 1;
• Each line contains q + 1 points, and each point lies in q + 1 lines;
• Any two points determine a unique line, and any two lines determine a unique point.
The Fano plane is thus a projective plane of order 2. More generally, if Fq is any finite field, then one can
define a projective plane P2q whose points and lines are the 1- and 2-dimensional subspaces F3q , respectively.
Note that the number of lines is the number of nonzero vectors up to scalar multiplication, hence (q 3 −
1)/(q − 1) = q 2 + q + 1.
A notorious open question is whether any other finite projective planes exist. The best general result known
is the Bruck–Ryser–Chowla theorem (1949), which states that if q ≡ 1 or 2 (mod 4), then q must be the sum
of two squares. In particular, there exists no projective plane of order 6. Order 10 is also known to be
impossible thanks to computer calculation, but the problem is open for other non-prime-power orders. It
is also open whether there exists a projective plane of prime-power order that is not isomorphic to P2q . One
readily available survey of the subject is by Perrott [Per16]. ◀
Representability can be tricky. As we have seen, U2 (4) can be represented over any field other than F2 ,
while the Fano plane is representable only over fields of characteristic 2. The point configuration below is
an affine representation of a rank-3 matroid over R, but the matroid is not representable over Q [Grü03,
pp. 93–94]. Put simply, it is impossible to construct a set of points with rational coordinates and exactly
these collinearities.
A regular matroid is one that is representable over every field. (For instance, we will see that graphic
matroids are regular.) For some matroids, the choice of field matters. For example, every uniform matroid
is representable over every infinite field, but as we have seen before, Uk (n) can be represented over Fq only
if n ≤ (q k − 1)/(q − 1). (For example, U2 (4) is not representable over F2 .) However, this inequality does not
suffice for representability; as mentioned above, the Fano plane cannot be represented over, say, F101 .
Recall that a minor of a matrix is the determinant of some square submatrix of M . A matrix is called totally
unimodular if every minor is either 0, 1, or −1.
3 But not self-dual as a matroid in the sense to be defined in §3.7.
75
Theorem 3.5.4. A matroid M is regular if and only if it can be represented by the columns of a totally unimodular
matrix.
One direction is easy: if M has a unimodular representation then the coefficients can be interpreted as lying
in any field, and the linear dependence of a set of columns does not depend on the choice of field (because
−1 ̸= 0 and 1 ̸= 0 in every field). The reverse direction is harder (see [Oxl92, chapter 6]), and the proof is
omitted. In fact, something more is true: M is regular if and only if it is binary (representable over F2 ) and
representable over at least one field of characteristic ̸= 2.
Theorem 3.5.5. Graphic matroids are regular.
Proof. Let G = (V, E) be a graph on vertex set V = [n], and let M = M (G) be the corresponding graphic
matroid. We can represent M by the matrix X whose columns are the vectors ei − ej for ij ∈ E. (Or ej − ei ;
it doesn’t matter, since scaling a vector does not change the matroid.) Here {e1 , . . . , en } is the standard
basis for Rn .
Consider any square submatrix XW B of X with rows W ⊆ V and columns B ⊆ A, where |W | = |B| = k >
0. If B contains a cycle v1 , . . . , vk then the columns are linearly dependent, because
so det XW B = 0. On the other hand, if B is acyclic, then I claim that det XW B ∈ {0, ±1}, which we will
prove by induction on k. The base case k = 1 follows because all entries of X are 0 or ±1. For k > 1, if
there is some vertex of W with no incident edge in B, then the corresponding row of XW B is zero and the
determinant vanishes. Otherwise, by the handshaking theorem, there must be some vertex w ∈ W incident
to exactly one edge b ∈ B. The corresponding row of XW B will have one entry ±1 and the rest zero.
Expanding on that row gives det XW B = ± detW \w,B\b , and we are done by induction. The same argument
shows that any set of columns corresponding to an acyclic edge set will in fact be linearly independent.
Example 3.5.6. The matrix
1 0 1 1
0 1 1 −1
represents U2 (4) over any field of characteristic ̸= 2, but the last two columns are dependent (in fact equal)
in characteristic 2. ◀
Example 3.5.7. There exist matroids that are not representable over any field. The smallest ones have
ground sets of size 8; one of these is the rank-4 Vámos matroid V8 [Oxl92, p. 511]. The smallest rank-3
example is the non-Pappus matroid.
Pappus’ Theorem from Euclidean geometry says that if a, b, c, A, B, C are distinct points in R2 such that
a, b, c and A, B, C are collinear, then x, y, z are collinear, where
76
c
b
a
y z
x
A
B
C
Accordingly, there is a rank-3 simple matroid on ground set E = {a, b, c, A, B, C, x, y, z} whose flats are
It turns out that deleting xyz from this list produces the family of closed sets of a matroid, called the non-
Pappus matroid NP. Since Pappus’ theorem can be proven using analytic geometry, and the equations that
say that x, y, z are collinear are valid over any field (i.e., involve only ±1 coefficients), it follows that NP is
not representable over any field. ◀
There are several ways to construct new matroids from old ones. We will begin with a boring but useful
one (direct sum) and then move on to the more exciting constructions of duality and deletion/contraction.
Definition 3.6.1. Let M1 , M2 be matroids on disjoint sets E1 , E2 , with basis systems B1 , B2 . The direct sum
M1 ⊕ M2 is the matroid on E1 ∪ E2 with basis system
B = {B1 ∪ B2 : B1 ∈ B1 , B2 ∈ B2 }.
If M1 , M2 are linear matroids whose ground sets span vector spaces V1 , V2 respectively, then M1 ⊕ M2 is the
matroid you get by regarding the vectors as living in V1 ⊕ V2 : the linear relations have to come either from
V1 or from V2 .
77
G1 G2 G
A useful corollary is that every graphic matroid arises from a connected graph. Actually, there may be many
different connected graphs that give rise to the same matroid, since in the previous construction it did not
matter which vertices of G1 and G2 were identified. This raises an interesting question: when does the
isomorphism type of a graphic matroid M (G) determine the graph G up to isomorphism?
Definition 3.6.2. A matroid that cannot be written nontrivially as a direct sum of two smaller matroids is
called connected or indecomposable.4
Proposition 3.6.3. Let G = (V, E) be a loopless graph. Then M (G) is indecomposable if and only if G is 2-
connected — i.e., not only is it connected, but so is every subgraph obtained by deleting a single vertex.
The “only if” direction is immediate: the discussion above implies that
M
M (G) = M (H)
H
In light of Proposition 3.6.3, it is natural to suspect that every 2-connected graph is determined up to iso-
morphism by its graphic matroid, but even this is not true; the two 2-connected graphs below are not
isomorphic, but have isomorphic graphic matroids.
As you should expect from an operation called “direct sum,” properties of M1 ⊕ M2 should be easily de-
ducible from those of its summands. In particular, direct sum is easy to describe in terms of the other
matroid axiomatizations we have studied. It is additive on rank functions: if A1 ⊆ E1 and A2 ⊆ E2 , then
rM1 ⊕M2 (A1 ∪ A2 ) = rM1 (A1 ) + rM2 (A2 ).
4 The first term is more common among matroid theorists, but I prefer “indecomposable” to avoid potential confusion with the
78
Similarly, the closure operator is A1 ∪ A2 = A1 ∪ A2 . The circuit system of the direct sum is just the (neces-
sarily disjoint) union of the circuit systems of the summands. Finally, the geometric lattice of a direct sum
is just the poset product of the lattices of the summands, i.e.,
L(M1 ⊕ M2 ) ∼
= L(M1 ) × L(M2 ),
subject to the order relations (F1 , F2 ) ≤ (F1′ , F2′ ) iff Fi ≤ Fi′ in L(Mi ) for each i.
3.7 Duality
Definition 3.7.1. Let M be a matroid on ground set |E| with basis system B. The dual matroid of M (also
known as the orthogonal matroid) is the matroid M ∗ on E with basis system
B ∗ = {E \ B : B ∈ B}.
We often write e∗ for elements of the ground set when talking about their behavior in the dual matroid.
Clearly the elements of B ∗ all have cardinality |E| − r(M ) (where r is the rank), and complementation
swaps the basis exchange conditions (B2) and (B2′ ), so if you believe that those conditions are logically
equivalent (Problem 3.2) then you also believe that B ∗ is a matroid basis system.
It is immediate from the definition that (M ∗ )∗ = M . In addition, the independent sets of M are the com-
plements of the spanning sets of M ∗ (since A ⊆ B for some B ∈ B if and only if E \ A ⊇ E \ B), and vice
versa. The rank function r∗ of the dual is given by
The dual of a vector matroid has an explicit description. Let E = {v1 , . . . , vn } ⊆ kr , and let M = M (E). We
may as well assume that E spans kr , so r ≤ n, and the representing matrix X = [v1 | · · · |vn ] ∈ kr×n has full
row rank r.
Let Y be any (n−r)×n matrix with rowspace(Y ) = nullspace(X). That is, the rows of Y span the orthogonal
complement of rowspace(X) with respect to the standard inner product.
Theorem 3.7.2. With this setup, the columns of Y are a representation for M ∗ .
Before proving this theorem, we will do an example that will make things clearer.
Example 3.7.3. Let E = {v1 , . . . , v5 } be the set of column vectors of the following matrix (over R, say):
1 0 0 2 1
X = 0 1 0 2 1 .
0 0 1 0 0
Notice that X has full row rank (it’s in row-echelon form, after all), so it represents a matroid of rank 3 on
5 elements. We could take Y to be the matrix
0 0 0 1 −2
Y = .
1 1 0 0 −1
79
Then Y has rank 2. Call its columns {v1∗ , . . . , v5∗ }; then the column bases are
{v1∗ , v4∗ }, {v1∗ , v5∗ }, {v2∗ , v4∗ }, {v2∗ , v5∗ }, {v4∗ , v5∗ },
whose (unstarred) complements (e.g., {v2 , v3 , v5 }, etc.) are precisely the column bases for X. In particular,
every basis of M contains v3 (so v3 is a coloop), which corresponds to the fact that no basis of M ∗ contains v3∗
(so v3∗ is a loop). This makes sense linear-algebraically: v3 is linearly independent of all the columns, so no
vector with a nonzero entry in the 3rd position is orthogonal to any row of M , so v3∗ is the zero vector. ◀
Proof of Theorem 3.7.2. First, note that invertible row operations on a matrix X ∈ kr×n (i.e., multiplication
on the left by an element of GLr (k)) do not change the matroid represented by its columns; they simply
change the basis of kr .
Let B be a basis of M , and reindex so that B = {v1 , . . . , vr }. We can then perform invertible row-operations
to put X into reduced row-echelon form, i.e.,
X = [Ir | A]
where Ir is the r × r identity matrix and A is arbitrary. It is easy to check that nullspace X = (rowspace X ∗ )T ,
where
X ∗ = [−AT | In−r ],
(this is a standard recipe). But then the last n − r elements of X ∗ , i.e., E ∗ \ B ∗ , are clearly a column basis. By
the same logic, every basis of X is the complement of a column basis of Y , and the converse is true because
X can be obtained from X ∗ in the same way that X ∗ is obtained from X. Therefore the columns of X and
X ∗ represent dual matroids. Meanwhile, any matrix Y with the same rowspace as X ∗ can be obtained from
it by invertible row operations, hence represents the same matroid.
Duality and graphic matroids. Let G be a connected planar graph, i.e., one that can be drawn in the plane
with no crossing edges. The planar dual is the graph G∗ whose vertices are the regions into which G divides
the plane, with two vertices of G∗ joined by an edge e∗ if the corresponding faces of G are separated by an
edge e of G. (So e∗ is drawn across e in the construction.)
f* G*
G
f
e*
e
80
If G is not planar then in fact M (G)∗ is not a graphic matroid (although it is certainly regular).
Definition 3.7.4. Let M be a matroid on E. A loop is an element of E that does not belongs to any basis of
M . A coloop is an element of E that belongs to every basis of M . An element of E that is neither a loop nor
a coloop is called ordinary (probably not standard terminology, but natural and useful).
In a linear matroid, a loop is a copy of the zero vector, while a coloop is a vector that is not in the span of all
the other vectors.
A cocircuit of M is by definition a circuit of the dual matroid M ∗ . A matroid can be described by its cocircuit
system, which satisfy the same axioms as those for circuits (Definition 3.4.6). Set-theoretically, a cocircuit
is a minimal set not contained in any basis of M ∗ , so it is a minimal set that intersects every basis of M
nontrivially. For a connected graph G, the cocircuits of the graphic matroid M (G) are the bonds of G: the
minimal edge sets K such that G − K is not connected. Every bond C ∗ is of the following form: there is a
partition V (G) = X ∪· Y such that C ∗ is the set of edges with one endpoint in each of X and Y , and both
G|X and G|Y are connected.
We can also describe deletion and contraction on the level of basis systems:
(
{B ∈ B(M ) : e ̸∈ B} if e is not a coloop,
B(M \e) =
{B − e : B ∈ B(M )} if e is a coloop,
(
{B − e : B ∈ B(M ), e ∈ B} if e is not a loop,
B(M/e) =
{B : B ∈ B(M )} if e is a loop.
Again, the terms come from graph theory. Deleting an edge e of a graph G means removing it from the
graph, while contracting an edge means to shrink it down so that its two endpoints merge into one. The
resulting graphs are called G\e and G/e, and these operations are consistent with the effect on graphic
matroids, i.e.,
M (G\e) = M (G)\e, M (G/e) = M (G)/e. (3.7)
y y
v e z v z v y z
w x
x x
w w
G G−e G/e
81
Notice that contracting can cause some edges to become parallel, and can cause other edges (namely, those
parallel to the edge being contracted) to become loops. In matroid language, deleting an element from a
simple matroid always yields a simple matroid, but the same is not true for contraction.
We can define deletion and contraction of sets as well as single elements. To delete (resp., contract) a set,
simply delete (resp., contract) each of its elements in some order.
1. For each A ⊆ E, the deletion M \A and contraction M/A are well-defined (i.e., do not depend on the order in
which elements of A are deleted or contracted).
2. In particular
I (M \A) = {I ⊆ E \ A : I ∈ I (M )},
I (M/A) = {I ⊆ E \ A : I ∪ B ∈ I (M )}
(M \e)∗ ∼
= M ∗ /e∗ and (M/e)∗ ∼
= M ∗ \e∗ .
Here is what deletion and contraction mean for vector matroids. Let V be a vector space over a field k, let
E ⊆ V be a set of vectors spanning V , let M = M (E), and let e ∈ E. Then:
1. M \e = M (E − e). (If we want to preserve the condition that the ground set spans the ambient space,
then e must not be a coloop.)
2. M/e is the matroid represented by the images of E − e in the quotient space V /ke. (Note that if e is a
loop then this quotient space is just V itself.)
Thus both operations preserve representability over k. For instance, to find an explicit representation
of M/e, apply a change of basis to V so that e is the ith standard basis vector, then simply erase the ith
coordinate of every vector in E − e.
Any matroid M ′ obtained from M by some sequence of deletions and contractions is called a minor of M .
Proposition 3.8.3. Every minor of a graphic (resp., linear, uniform) matroid is graphic (resp., linear, uniform).
Proof. The graphic case follows from (3.7), and the linear case from the previous discussion. For uniform
matroids, the definitions imply that
Uk (n)\e ∼
= Uk (n − 1) and Uk (n)/e ∼
= Uk−1 (n − 1)
Many invariants of matroids can be expressed recursively in terms of deletion and contraction. The follow-
ing fact is immediate from Definition 3.8.1.
82
Proposition 3.8.4. Let M be a matroid on ground set E, and let b(M ) denote the number of bases of M . Let e ∈ E;
then
b(M \e)
if e is a loop;
b(M ) = b(M/e) if e is a coloop;
b(M \e) + b(M/e) otherwise.
Example 3.8.5. If M ∼ Uk (n), then b(M ) = nk , and the recurrence of Proposition 3.8.4 is just the Pascal
=
relation nk = n−1 + n−1
k k−1 . ◀
Many other matroid invariants satisfy analogous recurrences involving deletion and contraction. In fact,
Proposition 3.8.4 is the tip of an iceberg that we will explore in Chapter 4.
3.9 Exercises
Problem 3.1. Determine, with proof, all pairs of integers k ≤ n such that there exists a graph G with
M (G) ∼= Uk (n). (Here Uk (n) denotes the uniform matroid of rank k on n elements; see Example 3.2.5.).
Hint: Use Proposition 3.8.3.
Solution: Uk (n) is graphic if and only if k ∈ {0, 1, n − 1, n}. For the “if” direction, we can construct a graph
G with M (G) ∼ = Uk (n) as follows:
There are several ways to show that these are the only possibilities. Here is my favorite. If M is a graphic
matroid, then so are M − e and M/e for every e in its ground set; by induction, so is any minor of M — i.e.,
any matroid formed from M by some sequence of deletions and contractions. The matroid U2 (4) is a minor
of any matroid Uk (n) with 2 ≤ k ≤ n − 2 (perform k − 2 contractions and n − 4 deletions), so it suffices to
show that U2 (4) is not graphic. If it were, the corresponding graph G would have three vertices and four
edges, and every two edges would form a spanning tree — but this is impossible because if G is loopless,
then two edges must be parallel.
Problem 3.2. Prove the equivalence of the two forms of the basis exchange condition (B2) and (B2′ ). (Hint:
Examine |B \ B ′ |.)
Solution: Suppose that B is a family of subsets of E. For reference, the axioms for a basis system are as
follows: for all B, B ′ ∈ B,
(B1) |B| = |B ′ |;
(B2) For all e ∈ B \ B ′ , there exists e′ ∈ B ′ \ B such that B − e + e′ ∈ B;
(B2′ ) For all e ∈ B \ B ′ , there exists e′ ∈ B ′ \ B such that B ′ + e \ e′ ∈ B.
and our job is to show that, given (B1), either of (B2) or (B2′ ) implies the other.
For two bases B, B ′ ∈ B ′ , define d(B, B ′ ) = |B \ B ′ | = |B ′ \ B| (the second equality follows from (B1)).
Both directions of the proof will proceed by induction on d(B, B ′ ).
83
First, we prove that (B2) implies (B2′ ).
Base case: If d(B, B ′ ) = 1, say B \ B ′ = {e} and B ′ \ B = {e′ }, then B ′ − e′ + e = B ∈ B, proving (B2′ ).
Inductive step: Suppose that d(B, B ′ ) > 1. Fix e ∈ B \ B ′ and choose f ∈ (B \ B ′ ) − e; note that this set is
nonempty because d(B, B ′ ) > 1. By (B2), there exists f ′ ∈ B ′ \ B such that B ′′ = B − f + f ′ ∈ B. Then
d(B ′ , B ′′ ) = d(B, B ′ ) − 1, and e ∈ B ′′ \ B ′ . By induction, there exists e′ ∈ B ′ \ B ′′ such that B ′ − e′ + e ∈ B.
Moreover, e′ ∈ B ′ \ B ′′ ⊆ B ′ \ B, establishing (B2′ ).
Base case: If d(B, B ′ ) = 1, say B \ B ′ = {e} and B ′ \ B = {e′ }, then B ′ + e − e′ = B ∈ B, proving (B2).
Fix e ∈ B \ B ′ and choose f ∈ (B \ B ′ ) − e; note that this set is nonempty because d(B, B ′ ) > 1. By (B2′ ),
there exists f ′ ∈ B ′ \ B such that B ′′ = B ′ − f ′ + f ∈ B. Then d(B, B ′′ ) = d(B, B ′ ) − 1. By induction, there
exists e′ ∈ B ′′ \ B such that B − e + e′ ∈ B. But in fact e′ ∈ B ′ , because B ′′ \ B = (B ′ \ B) − f , and we have
established (B2).
Problem 3.3. Let A be a nonempty family of subsets of a finite set E. Prove that A is a matroid basis
system if and only if the following two conditions hold:
Solution: ( =⇒ ) Suppose that B is a matroid basis system, satisfying (B1) and (B2) of Definition 3.4.4. First,
(A1) is immediate from (B1).
( ⇐= ) Suppose that A is a set family satisfying (A1) and (A2). Suppose that not all members of A are the
same size. Let A1 , A2 ∈ A such that |A1 | > |A2 |, and such that m = |A2 \ A1 | is minimal among all such
pairs. If m = 0 then A2 ⊂ A1 , contradicting (A1). So m < 0, and letting B = A1 ∩ A2 , we can write
A1 = (A1 ∩ A2 ) ∪ {a1 , . . . , an },
A2 = (A1 ∩ A2 ) ∪ {b1 , . . . , bm }
where n > m. Define
X = (A1 ∩ A2 ) ∪ {a1 , . . . , am },
Y = (A1 ∩ A2 ) ∪ {a1 , . . . , am , b1 , . . . , bm } = A2 ∪ {a1 , . . . , am }
so that A1 ⊋ X ⊊ Y ⊋ A2 . By (A2), there exists A ∈ A such that X ⊊ A ⊊ Y (equality in either case would
contradict (A1)). Without loss of generality we may assume that
A = (A1 ∩ A2 ) ∪ {a1 , . . . , am } ∪ {b1 , . . . , bℓ }
where 0 < ℓ < m. But then |A| > |A2 | and |A2 \ A| = m − ℓ < m, contradicting the choice of the pair A1 , A2 .
It follows that all elements of A have the same size, establishing (B1).
84
Now, let A1 , A2 be distinct elements of A and let e ∈ A1 \ A2 . Let X = A1 \ {e} and Y = A1 ∪ A2 . By (A2)
there exists A ∈ A satisfying X ⊆ A ⊆ Y . But since we already know (B1) holds, we must have |A| = |A1 |
and therefore A = A1 \ {e} ∪ {e′ } for some e′ ∈ A2 \ A1 .
( ⇐= ) Suppose that A is a set family satisfying (A1) and (A2). Let A1 , A2 ∈ A and e ∈ A1 \ A2 . Let
X = A1 − e and Y = A2 ∪ X, so that the hypotheses of (A2) hold. Then (A1) implies that X itself is not a
basis, so A must be of the form A1 − e + e′ for some e′ ∈ Y \ A1 = A2 \ A1 , establishing (B2).
To prove (B1), observe that if |A1 | > |A2 |, then replacing A1 with A1 − e + e′ and repeating this process
eventually produces a basis that contains A2 and has the same size as A1 , which violates (A1).
Problem 3.4. (Proposed by Kevin Adams.) Let B, B ′ be bases of a matroid M . Prove that there exists a
bijection ϕ : B \ B ′ → B ′ \ B such that B − e + ϕ(e) is a basis of M for every e ∈ B \ B ′ .
Solution: Let X = B \ B ′ and Y = B ′ \ B. Consider the bipartite graph G with vertex set B△B ′ = X ∪ Y ,
with an edge between x ∈ X and y ∈ Y if B − x + y is a basis. Then the problem is to show that G has a
perfect matching, so that we can define ϕ to be the G-spouse of an element of B \ B ′ . By Hall’s Theorem, it
is sufficient to show that |N (V )| ≥ |V | for all V ⊆ X, where N denotes neighborhood in G.
Let V ⊆ X. Then for every x ∈ V and y ∈ Y \ N (V ), there is a circuit Cxy ⊆ B − x + y. Each such circuit
certainly contains w, since B \ v is independent. If Cxy ̸= Cx′ y for some x, x′ ∈ V and y ∈ Y \ N (V ),
then the circuit exchange axiom says that Cxy ∪ Cx′ y − y contains a circuit, but that is impossible because
Cxy ∪ Cx′ y − y ⊆ B. Therefore Cxy depends only on y, and we can safely change the notation to Cy . But
then \
Cy ∈ (X − x + y) = (X \ V ) + y
x∈V
which says that every element in Y \N (V ) is in the closure of X \V , and in particular r(Y \N (V )) ≤ r(X \V ).
But these are both independent sets, so it follows that |Y \ N (V )| ≤ |X \ V |, hence |N (V )| ≥ |V |, which is
Hall’s criterion. Note that the bijection ϕ−1 has the dual property.
Problem 3.5. Prove Proposition 3.4.5, which describes the cryptomorphism between matroid independence
systems and matroid basis systems.
Solution: (1) Let I be a matroid independence system on E and let B be the family of maximal elements
of E. Then B is nonempty because I is nonempty, and the donation condition implies that the maximal
elements of B all have the same cardinality. To establish the basis exchange condition: Let B, B ′ be distinct
elements of B. Let e ∈ B \ B ′ and I = B − e. Then I ∈ I and |I| = |B ′ | − 1, so by the donation axiom, there
exists an element e′ ∈ B ′ \ I such that I + e′ ∈ B. On the other hand, B ′ \ I = B ′ \ B and I + e′ = B \ x + e,
so we have established (B2).
(2) Let B be a matroid basis system and let I = B∈B 2B . By definition this is a simplicial complex, so we
S
just need to show that the donation axiom holds. Therefore, let I, J ∈ I with |I| < |J|. Let BI and BJ be
elements of B containing I and J, respectively. Let B0 = BI . If B0 \ I contains an element e ̸∈ BJ , then
choose e′ ∈ BJ so that B1 = B0 − e + e′ ∈ B. Thus I ⊆ B1 and |(B1 \ I) \ Bj | = |(B0 \ I) \ Bj | − 1. Repeat
this process a total of |(B0 \ I) \ Bj | times until we obtain a basis B ′′ = I ∪ K, with K ⊆ BJ . But J, K ⊆ BJ
and |J| + |K| = |J| + |BJ | − |I| > |BJ | (since |J| > |I|). It follows that J ∩ K ̸= 0. Let x be an element of
J ∩ K; then I + x ⊆ BI , so I + x ∈ I and we have established the donation axiom.
Problem 3.6. Prove Proposition 3.4.7, which describes the cryptomorphism between matroid independence
systems and matroid circuit systems. (Hint: The hardest part is showing that if C is a matroid circuit system
then the family I of sets containing no circuit satisfies (I3). Under the assumption that (I3) fails for some
pair I, J with |I| < |J|, use circuit exchange to build a sequence of collections of circuits in I ∪ J that avoid
more and more elements of I, eventually producing a circuit in J and thus producing a contradiction.)
85
Solution: Part 1: Let I be a matroid independence system on E and define
C = {C ⊆ E : C ̸∈ I and C ′ ∈ I ∀C ′ ⊊ C}.
Then (I1) directly implies (C1). Moreover, the elements of C are pairwise incomparable, because if C, C ′ ∈
C and C ⊊ C ′ , then both C ∈ I and C ̸∈ I , which is silly; this establishes (C2).
Now we prove (C3). Let C, C ′ be distinct circuits. Then C ∩ C ′ is a proper subset of both C, C ′ , hence is
independent. Let x ∈ C ∩ C ′ and let D = C ∪ C ′ − x. Note that D \ (C ∩ C ′ ) = (C \ C ′ ) ∪ (C ′ \ C) = C△C ′ .
Suppose that D is independent. By (I3), D can repeatedly donate elements to C ∩C ′ to produce independent
sets
(C \ C ′ ) + y1 , (C \ C ′ ) + y1 + y2 , . . . , Z = (C ∩ C ′ ) ∪ D − y
where either y ∈ C ∩ C ′ or y ∈ C ′ ∩ C. (Note that all sets in this list, other than Z, have smaller cardinality
than D, which is why we can apply donation repeatedly.) But then either Z ⊇ C ′ or Z ⊇ C, which are both
impossible by the definition of C .
I = {I ⊆ E : C ̸⊆ I ∀C ∈ C }.
Evidently I satisfies (I1) and (I2); the substantial part is proving (I3). Accordingly, suppose I, J ∈ I with
|I| < |J|. We wish to show that there exists x ∈ J \ I such that I ∪ x ∈ I . If I ⊆ J then this is immediate,
so suppose I ̸⊆ J.
Let I \ J = {x1 , . . . , xq } and J \ I = {y1 , . . . , yp }; note that q < p. For each i ∈ [p] there is a circuit
Ci ⊆ I ∪ {yi }. Note that yi ∈ Ci and that Ci is uniquely determined — if there are two such Ci , then circuit
exchange produces a circuit in I, which is a contradiction. Consider the following algorithm.
for r from 1 to q:
If xr ̸∈ Ci for all i ∈ [q + 1]:
do nothing.
Else:
Reindex yr+1 , . . . , yq−r+2 and Cr+1 , . . . , Cq−r+2 so that xr ∈ Cq−r+2 .
For i from 1 to q − r + 1:
(A) if xr ̸∈ Ci , then let Ci′ = Ci .
(B) Else, let Ci′ be a circuit in Ci ∪ Cq+1 \ {xr } ⊆ I + yi + yq+1 .
I claim that yi ∈ Ci′ for every i ∈ {1, 2, . . . , q − r + 1}. (Note that this set is nonempty because r ≤ q.) If the
claim is false, then Ci′ ⊆ I ∪ yq+1 which means that Ci′ = Cq+1 (the only circuit in I ∪ yq+1 !), so we must have
chosen option (B) rather than (A), which means that xr ̸∈ Ci′ , but our choice of reindexing says xr ∈ Cq+1 ,
which is a contradiction.
Part 3: We check that these constructions are inverses. Let I be a matroid independence system and let
If A ∈ I , then every subset of A is also a member of I (by (I2)), hence none of them is a member of C , so
A ∈ I ′ . Conversely, if A ̸∈ I ′ , then A must have some subset A′ ∈ C . Then A′ ̸∈ I , so A ̸∈ I as well.
86
Meanwhile, let C be a circuit system and let
If C ∈ C ′ , then C ̸∈ I but every proper subset of C is a member of I . Therefore C has some subset
C ′′ ∈ C , but C ′′ is not a subset of any proper subset of C; it follows that C ′′ = C ∈ C . Hence C ′ ⊆ C , and
this argument can be reversed to show C ⊆ C ′ .
Problem 3.7. Let M be a matroid on ground set E. Suppose thereL is a partition of E into disjoint sets
n
E1 , . . . , En such that r(E) = r(E1 ) + · · · + r(Ek ). Prove that M = i=1 Mi , where Mi = M |Ei . (Note:
This fact provides an algorithm, albeit not necessarily an efficient one, for testing whether a matroid is
connected.)
Solution: We will show that the bases of M are precisely the sets B1 ∪ · · · ∪ Bn , where Bi is a basis of Mi for
every i.
where the first inequality comes from submodularity. But r(B) = r(M ), so equality holds throughout, and
in particular it holds for each summand. Therefore, for each i, the set Bi spans Mi , and it is independent
(by virtue of being a subset of the independent set B), so it is a basis of Mi .
B̄ = B1 ∪ · · · ∪ Bn ⊃ B̄1 ∪ · · · ∪ B̄n = E1 ∪ · · · ∪ En = E
P P
so B is a spanning set, and its cardinality is i |Bi | = i r(Mi ) = r(M ), so it is a basis.
Problem 3.8. Let M be a matroid on ground set E with rank function r : 2E → N. Prove that the rank
function r∗ of the dual matroid M ∗ is given by r∗ (A) = r(E \ A) + |A| − r(E) for all A ⊆ E.
Solution: Let B and B ∗ be the basis systems of M and M ∗ respectively, and let A ⊆ E. Here we go:
Alternative solution method, found in 2022 by Kathryn Cole, Alireza Masaeli, and Dania Morales: Show that r∗ is
a matroid rank function, i.e., it satisfies the conditions of Definition 3.2.4. Then show that the bases of the
matroid thus constructed are precisely the complements of bases of M .
Problem 3.9. Let M be a matroid on E. A set S ⊆ E is called spanning if it contains a basis. Let S be the
set of all spanning sets.
87
(a) Express S in terms of (i) the rank function r of M ; (ii) its closure operator A 7→ Ā; (iii) its lattice of
flats L. (You don’t have to prove anything — just give the construction.)
(b) Formulate axioms that could be used to define a matroid via its system of spanning sets. (Hint:
Describe spanning sets in terms of the dual matroid M ∗ .)
Solution: (a)
S = {S ⊆ E : S ⊇ B for some B ∈ B}
= {S ⊆ E : r(S) = r(E)}
= {S ⊆ E : S̄ = E}
_
= {S ⊆ E : s = 1̂L }.
s∈S
Therefore, we can translate the axioms for I (M ∗ ) to be an independence system into axioms for S to be a
spanning system:
(S1) E ∈ S ;
(S2) if S ∈ S and T ⊇ S, then T ∈ S ;
(S3) if S, T ∈ S with |S| < |T |, then there exists e ∈ T \ S such that T \e ∈ S .
on E. Let w : E → R≥0
Problem 3.10. Let E be a finite set and let ∆ be an abstract simplicial complex P
be any function; think of w(e) as the “weight” of e. For A ⊆ E, define w(A) = e∈A w(e). Consider the
problem of maximizing w(A) over all facets A. A naive approach is the following greedy algorithm:
Step 1: Let A = ∅.
Step 2: If A is a facet of ∆, stop.
Otherwise, find e ∈ E \ A of maximal weight such that A + e ∈ ∆
(if there are several such e, pick one at random), and replace A with A + e.
Step 3: Repeat Step 2 until A is a facet of ∆.
This algorithm may or may not work for a given ∆ and w. Prove the following facts:
(a) Construct a simplicial complex and a weight function for which this algorithm does not produce a
facet of maximal weight. (Hint: The smallest example has |E| = 3.)
(b) Prove that the following two conditions are equivalent:
(i) The greedy algorithm produces a facet of maximal weight for every weight function w.
(ii) ∆ is a matroid independence system.
Note: This result does follow from Theorem 6.5.1. However, that is a substantial result, so don’t use it
unless you first do Problem 6.9. It is possible to do this exercise by working directly with the definition of
a matroid independence system.
Solution: (1) Consider the simplicial complex ∆ = ⟨a, bc⟩ with weight function w(a) = 3, w(b) = w(c) = 2.
88
(2) ( ⇐= ) Suppose that ∆ is a matroid independence system and let w be a weight function. Let B be a
basis of maximum weight and let B0 be the output of the greedy algorithm. If B = B0 then we are done;
otherwise, let e be an element of B \B0 . By basis exchange, there is some f ∈ B0 \B such that B1 = B0 −f +e
is a basis. I claim that the greedy algorithm must have considered f before e — otherwise, at the stage when
e was considered, the set A would have been a subset of B0 − f and e would have been added, which it
wasn’t. But this means that w(f ) ≥ w(e), so w(B1 ) ≤ w(B0 ). Also, |B1 ∩ B| = |B0 ∩ B| + 1 (since B1 was
formed from B0 by replacing a non-element of B with an element of B). Repeating this process, we obtain
a sequence of bases B0 , B1 , . . . with |Bk ∩ B| = |B0 ∩ B| + k and w(B0 ) ≥ w(B1 ) ≥ · · · . Eventually we get
Bℓ = B, but since B was assumed maximum weight, equality must hold throughout. In particular, B0 has
maximum weight.
( =⇒ ) Suppose the greedy algorithm always works. Let I and J be independent sets with |I| < |J|. Let
p = |I|, q = |J|, and ϵ a very small positive number (it suffices to take ϵ < (q − p)/p). Define a weight
function
1 + ϵ for x ∈ I,
w(x) = 1 for x ∈ J \ I,
0 otherwise.
The greedy algorithm begins by selecting all elements of I, whose weight is p(1+ϵ). But I is not a maximum-
weight face because p(1 + ϵ) < q = w(J). Therefore, the algorithm must add at least one element of J \ I at
some point, which is precisely the statement that (I3) holds.
Problem 3.11. Prove Proposition 3.8.2.
so (M − e)∗ = M ∗ /e∗ . Dualizing both sides gives M − e = (M ∗ /e∗ )∗ , and then switching the roles of M
and M ∗ we get M ∗ − e∗ = (M/e)∗ .
Problem 3.12. Let X and Y be disjoint sets of vertices, and let B be an X, Y -bipartite graph: that is, every
edge of B has one endpoint in each of X and Y . For V = {x1 , . . . , xn } ⊆ X, a transversal of V is a set
W = {y1 , . . . , yn } ⊆ Y such that xi yi is an edge of B. (The set of all edges xi yi is called a matching.) Let I
be the family of all subsets of X that have a transversal; in particular I is a simplicial complex.
Prove that I is in fact a matroid independence system by verifying that the donation condition (I3) holds.
(Suggestion: Write down an example or two of a pair of independent sets I, J with |I| < |J|, and use the
corresponding matchings to find a systematic way of choosing a vertex that J can donate to I.) These ma-
troids are called transversal matroids; along with linear and graphic matroids, they are the other “classical”
examples of matroids in combinatorics.
Solution: Let I and J be independent sets with |I| < |J| and let MI and MJ be the corresponding matchings.
Consider the graph with edges MI ∪ MJ . No vertex has more than two edges in this graph, which means
that every component must either be a path or a cycle, with edges alternating between MI and MJ (since
no vertex has more than one edge in either of them). In particular, since |I| < |J|, at least one component is
a path of the form
x1 J y2 I x1 J y2 I x3 ··· yn−1 I xn J yn
89
with x1 ̸∈ I. But then I + x1 has a transversal, whose matching is obtained by replacing the n − 1 edges
yi xi+1 with the n edges xi yi .
Problem 3.13. Fix positive integers n ≥ r.
Solution: We just prove (b), which as we have noticed implies (a). First, we claim that the condition
|S△S ′ | ≥ 2j + 4 is equivalent to |S ∩ S ′ | ≤ k − 2. Observe that
|S ∪ S ′ | + |S ∩ S ′ | = |S| + |S ′ | = 2k + 2j,
|S ∪ S ′ | − |S ∩ S ′ | = |S△S ′ | ≥ 2j + 4,
so subtracting gives
2|S ∩ S ′ | ≤ (2k + 2j) − (2j + 4) = 2k − 4
and dividing by 2 proves the claim.
Case I: Two of the Si , say S1 and S2 , are distinct. By the Claim we have |S1 ∩ S2 | ≤ k − 2 by the Claim. On
the other hand, S1 ∩ S2 ⊇ C1 ∩ C2 = B \ e, and |B \ e| = k − 1, which is a contradiction.
S1 ⊇ C1 ∪ · · · ∪ Cr = (B \ e) ∪ {f1 , . . . , fr } ⊇ (B ∩ B ′ ) ∪ (B ′ \ B) = B ′
Solution: If r = 1, then relaxation simply changes a loop to a nonloop. So assume r > 1. Certainly every
two elements of B satisfy basis exchange, so we need only to check basis exchange between an old basis B
and a new basis A.
First, suppose a ∈ A \ B. Note that a is not a loop (since A is a circuit of size > 1), so there exists some basis
B ′ ∈ B containing a. By basis exchange for B. there exists some e ∈ B such that B \ e ∪ a is a basis.
Second, suppose e ∈ B \ A. Since A is a flat of B, we must have r(A ∪ e) ≥ r(A) + 1 = r, and in addition
|A ∪ e| = r + 1. Therefore A ∪ e must contain a basis B ′ ∈ B, which cannot be A, hence must be of the form
A ∪ e \ a for some a ∈ A, as desired.
90
Problem 3.15. (Requires a bit of abstract algebra.) Let n be a positive integer, and let ζ be a primitive nth
root of unity. The cyclotomic matroid Yn is represented over Q by the numbers 1, ζ, ζ 2 , . . . , ζ n−1 , regarded as
elements of the cyclotomic field extension Q(ζ). Thus, the rank of Yn is the dimension of Q(ζ) as a Q-vector
space, which is given by the Euler ϕ function. Prove the following:
This problem is near and dear to my heart; the answer (more generally, a characterization of Yn for all n)
appears in [MR05].
E0 = {ζ 0 , ζ q , ζ 2q , . . . , ζ (m−1)q },
E1 = {ζ 1 , ζ q+1 , ζ 2q+1 , . . . , ζ (m−1)q+1 },
...
Eq−1 = {ζ q−1 , ζ 2q−1 , . . . , ζ n−1 },
and let Mi = Yn |Ei . Then E0 is just the set of qth roots of unity, so M0 ∼
= Yq . In fact Mi ∼
= Yq for each i, since
multiplying by the nonzero scalar ζ i defines a matroid isomorphism M0 → Mi .
Now, since ϕ(n) = qϕ(m), the desired result follows from Problem 3.7.
91
Chapter 4
Throughout this section, let M be a matroid on ground set E with rank function r, and let n = |E|.
Example 4.1.2. If E = ∅ then TM (x, y) = 1. Mildly less trivially, if every element is a coloop, then r(A) = |A|
for all A, so X
TM = (x − 1)n−|A| = (x − 1 + 1)n = xn
A⊆E
by the binomial theorem. If every element is a loop, then the rank function is identically zero and we get
X
TM = (y − 1)|A| = y n .
A⊆E
◀
Example 4.1.3. For uniform matroids, corank and nullity depend only on cardinality, making their Tutte
polynomials easy to compute. U1 (2) has one set with corank 1 and nullity 0 (the empty set), two singleton
sets with corank 0 and nullity 0, and one doubleton with corank 0 and nullity 1, so
TU1 (2) = (x − 1) + 2 + (y − 1) = x + y.
92
Similarly,
◀
Example 4.1.4. Let G be the graph below (known as the “diamond”):
a e c
Many invariants of M can be obtained by specializing the variables x, y appropriately. Some easy ones:
1. TM (2, 2) = A⊆E 1 = 2|E| . (Or, if you like, |E| = log2 TM (2, 2).)
P
2. Consider TM (1, 1). This kills off all summands whose corank is nonzero (i.e., all non-spanning sets)
and whose nullity is nonzero (i.e., all non-independent sets). What’s left are the bases, each of which
contributes a summand of 1. So TM (1, 1) = b(M ), the number of bases. We previously observed that
this quantity satisfies a deletion/contraction recurrence (Prop. 3.8.4); this will show up again soon.
3. Similarly, TM (1, 2) and TM (2, 1) count respectively the number of spanning sets and the number of
independent sets.
4. A little more generally, we can enumerate independent and spanning sets by their cardinality:
X
q |A| = q r(M ) T (1/q + 1, 1);
A⊆E independent
X
q |A| = q r(M ) T (1, 1/q + 1).
A⊆E spanning
5. TM (0, 1) is (up to a sign) the reduced Euler characteristic (see (6.2)) of the independence complex
93
of M :
X X
TM (0, 1) = (−1)r(E)−r(A) 0|A|−r(A) = (−1)r(E)−r(A)
A⊆E A⊆E independent
X
r(E) |A|
= (−1) (−1)
A∈I (M )
The fundamental theorem about the Tutte polynomial is that it satisfies a deletion/contraction recurrence.
In a sense it is the most general such invariant — we will give a “recipe theorem” that expresses any
deletion/contraction invariant as a Tutte polynomial specialization (more or less).
Theorem 4.1.5. The Tutte polynomial satisfies (and can be computed by) the following Tutte recurrence:
(T1) If E = ∅, then TM = 1.
(T2a) If e ∈ E is a loop, then TM = yTM \e .
(T2b) If e ∈ E is a coloop, then TM = xTM/e .
(T3) If e ∈ E is ordinary, then TM = TM \e + TM/e .
We can use this recurrence to compute the Tutte polynomial, by picking one element at a time to delete
and contract. The miracle is that it doesn’t matter what order we choose on the elements of E — all orders will
give the same final result! (In the case that M is a uniform matroid, then it is clear at this point that TM is
well-defined by the Tutte recurrence, because, up to isomorphism, M \e and M/e are independent of the
choices of e ∈ E.)
The Tutte recurrence says that we can represent a calculation of TM by a binary tree, with a branch for each
deletion/contraction:
M
M \e M \f
94
Example 4.1.8. Consider the diamond of Example 4.1.4. One possibility is to recurse on edge a (or equiva-
lently on b, c, or d). When we delete a, the edge d becomes a coloop, and contracting it produces a copy of
K3 . Therefore
T (G\a) = x(x2 + x + y)
by Example 4.1.7. Next, apply the Tutte recurrence to the edge b in G/a. The graph G/a\b has a coloop c,
contracting which produces a digon. Meanwhile, M (G/a/b) ∼ = U1 (3). Therefore
x(x2 + x + y)
x(x + y) x + y + y2
95
x3 x2 + x + y x(x + y) y(x + y)
Proof of Theorem 4.1.5. Let M be a matroid on ground set E, let e ∈ E, and let r′ and r′′ be the rank func-
tions of M \e and M/e respectively. The definitions of rank function, deletion, and contraction imply the
following, for A ⊆ E − e:
96
For (T2b), let e be a coloop. Then
X
TM = X r(E)−r(A) Y |A|−r(A)
A⊆E
X X
= X r(E)−r(A) Y |A|−r(A) + X r(E)−r(B) Y |B|−r(B)
e̸∈A⊆E e∈B⊆E
(r ′′ (E−e)+1)−r ′′ (A) |A|−r ′′ (A) ′′
(E−e)+1)−(r ′′ (A)+1) ′′
X X
= X Y + X (r Y |A|+1−(r (A)+1)
A⊆E−e A⊆E−e
r ′′ (E−e)+1−r ′′ (A) |A|−r ′′ (A) ′′
(E−e)−r ′′ (A) ′′
X X
= X Y + Xr Y |A|−r (A)
A⊆E−e A⊆E−e
r ′′ (E−e)−r ′′ (A) |A|−r ′′ (A)
X
= (X + 1) X Y = xTM/e .
A⊆E−e
A⊆E−e A⊆E−e
= TM \e + TM/e .
Some easy and useful observations (which illustrate, among other things, that both the rank-nullity and
recursive forms are valuable tools):
1. The Tutte polynomial is multiplicative on direct sums, i.e., TM1 ⊕M2 = TM1 TM2 . This is probably easier
to see from the rank-nullity generating function than from the recurrence.
2. Duality interchanges x and y, i.e.,
TM (x, y) = TM ∗ (y, x). (4.3)
The proof is left as an exercise (Problem 4.1). It can be deduced either from the Tutte recurrence (since
duality interchanges deletion and contraction; see Prop. (3.8.2)) or from the corank-nullity generating
function, by expressing r∗ in terms of r (see Problem 3.8).
3. The Tutte recurrence implies that every coefficient of TM is a nonnegative integer, a property which is
not obvious from the closed formula (4.1).
4.2 Recipes
The Tutte polynomial is often referred to as “the universal deletion/contraction invariant for matroids”:
every invariant that satisfies a deletion/contraction-type recurrence can be recovered from the Tutte poly-
nomial. This can be made completely explicit: the results in this section describe how to “reverse-engineer”
a general deletion/contraction recurrence for a graph or matroid isomorphism invariant to express it in
terms of the Tutte polynomial.
97
Theorem 4.2.1 (Tutte Recipe Theorem for Matroids). Let u(M ) be a matroid isomorphism invariant that satisfies
a recurrence of the form
1 if E = ∅,
Xu(M/e) if e ∈ E is a coloop,
u(M ) =
Y u(M \e) if e ∈ E is a loop,
au(M/e) + bu(M \e) if e ∈ E is ordinary
where E denotes the ground set of M and X, Y, a, b are either indeterminates or numbers, with a, b ̸= 0. Then
Proof. Denote by r(M ) and n(M ) the rank and nullity of M , respectively. Note that
whenever deletion and contraction are well-defined. Define a new matroid invariant
and rewrite the recurrence in terms of ũ, abbreviating r = r(M ) and n = n(M ), to obtain
1 if E = ∅,
Xar−1 bn ũ(M/e)
if e ∈ E is a coloop,
ar bn ũ(M ) =
Y ar bn−1 ũ(M \e) if e ∈ E is a loop,
r n
a b ũ(M/e) + ar bn ũ(M \e) if e ∈ E is ordinary.
Setting X = xa and Y = yb, we see that ũ(M ) = TM (x, y) = TM (X/a, Y /b) by Theorem 4.1.5, and rewriting
in terms of u(M ) gives the desired formula.
where G = (V, E) and X, Y, a, b, c are either indeterminates or numbers (with b, c ̸= 0). Then
We omit the proof, which is similar to that of the previous result. A couple of minor complications are that
many deletion/contraction graph invariants involve the numbers of vertices or components, which cannot
be deduced from the matroid of a graph. Also, while deletion and contraction of a cut-edge of a graph
produce two isomorphic matroids, they do not produce two isomorphic graphs (so, no, that’s not a misprint
in the coloop case of Theorem 4.2.2). The invariant U is described by Bollobás as “the universal form of the
Tutte polynomial.”
98
4.3 Basis activities
We know that TM (x, y) has nonnegative integer coefficients and that TM (1, 1) is the number of bases of
M . These observations suggest that we should be able to interpret the Tutte polynomial as a generating
function for bases: that is, there should be combinatorially defined functions i, e : B(M ) → N such that
X
TM (x, y) = xi(B) y e(B) .
B∈B(M )
In fact, this is the case. The tricky part is that i(B) and e(B) must be defined with respect to a total order
e1 < · · · < en on the ground set E, so they are not really invariants of B itself. However, another miracle
occurs: the Tutte polynomial itself is independent of the choice of total order.
Definition 4.3.1. Let M be a matroid on E with basis system B and let B ∈ B. For e ∈ B, the fundamental
cocircuit of e with respect to B, denoted C ∗ (e, B), is the unique cocircuit in (E \ B) + e. That is,
C ∗ (e, B) = {e′ : B − e + e′ ∈ B}.
Dually, for e ̸∈ B, then the fundamental circuit of e with respect to B, denoted C(e, B), is the unique circuit
in B + e. That is,
C(e, B) = {e′ : B + e − e′ ∈ B}.
In other words, the fundamental cocircuit consists of e together with all elements outside B that can re-
place e in a basis exchange, while the fundamental circuit consists of e together with all elements outside B
that can be replaced by e.
Suppose that M = M (G), where G its a connected graph, and B is a spanning tree. For all e ∈ B, the graph
B − e has two components, say X and Y , and C ∗ (e, B) is the set of all edges with one endpoint in each of
X and Y . Dually, if e ̸∈ B, then B + e has exactly one cycle, and that cycle is C(e, B).
If M is a vector matroid, then C ∗ (e, B) consists of all vectors not in the codimension-1 subspace spanned
by B − e, and C(e, B) is the unique linearly dependent subset of B + e.
Definition 4.3.2. Let M be a matroid on a totally ordered vertex set E = {e1 < · · · < en }, and let B be a
basis of M . An element e ∈ B is internally active with respect to B if e is the minimal element of C ∗ (e, B).
An element e ̸∈ B is externally active with respect to B if e is the minimal element of C(e, B). We set
i(B) = #{e ∈ B : e is internally active with respect to B}
= #{edges of B that cannot be replaced by anything smaller outside B},
e(B) = #{e ∈ E \ B : e is externally active with respect to B}
= #{edges of E \ B that cannot replaced anything smaller inside B}.
Note that these numbers depend on the choice of ordering of E.
Example 4.3.3. Let G be the graph with edges labeled as shown below, and let B be the spanning tree
{e2 , e4 , e5 } shown in red. The middle figure shows C(e1 , B), and the right-hand figure shows C ∗ (e5 , B).
e1 e1 e1
e2 e5 e4 e2 e5 e4 e2 e5 e4
e3 e3 e3
99
Here are some fundamental circuits and cocircuits:
C(e1 , B) = {e1 , e4 , e5 } so e1 is externally active;
C(e3 , B) = {e2 , e3 , e5 } so e3 is not externally active;
For instance, in Example 4.3.3, the spanning tree B contributes the monomial xy = x1 y 1 to T (G; x, y).
Tutte’s
P original paper [Tut54] actually defined the Tutte polynomial (which he called the “dichromate”)
as B∈B(M ) xi(B) y e(B) (rather than the corank-nullity generating function), then proved it that obeys the
deletion/contraction recurrence. Like the proof of Theorem 4.1.5, this result requires careful bookkeeping
but is not conceptually difficult. Note in particular that if e is a loop (resp. coloop), then e ̸∈ B (resp. e ∈ B)
for every basis B, and C(e, B) = {e} (resp. C ∗ (e, B) = {e}), so e is externally (resp. internally) active with
respect to B, so the generating function (4.4) is divisible by y (resp. x).
We first show that the characteristic polynomial of a geometric lattice is a specialization of the Tutte poly-
nomial of the corresponding matroid.
Theorem 4.4.1. Let M be a simple matroid on E with rank function r and lattice of flats !L. Then
χ(L; k) = (−1)r(M ) TM (1 − k, 0).
We now claim that f (K) = µL (0̂, K). For each flat K ∈ L, let
X
g(K) = f (J)
J∈L: J⊆K
100
so that by Möbius inversion (2.3a)
X
f (K) = µ(J, K)g(J). (4.5)
J∈L: J⊆K
Theorem 4.4.1 gives another proof that the Möbius function of a semimodular lattice L weakly alternates in
sign, or specifically that (−1)r(L) µ(L) ≥ 0 (Theorem 2.4.7). First, if L is not geometric, or equivalently not
atomic, then by Corollary 2.4.10 µ(L) = 0. Second, if L is geometric, then by (2.7) and Theorem 4.4.1
The characteristic polynomial of a graphic matroid has a classical combinatorial interpretation in terms of
colorings. Let G = (V, E) be a connected graph. Recall that a k-coloring of G is a function f : V → [k], and
a coloring is proper if f (v) ̸= f (w) whenever vertices v and w are adjacent. We showed in Example 2.3.5
that the function
pG (k) = number of proper k-colorings of G
is a polynomial in k, called the chromatic polynomial of G. In fact pG (k) = k · χK(G) (k). We can also prove
this fact via deletion/contraction.
• If G has a loop, then its endpoints automatically have the same color, so it’s impossible to color G
properly and pG (k) = 0.
• If G = Kn , then all vertices must have different colors. There are k choices for f (1), k − 1 choices for
f (2), etc., so pKn (k) = k(k − 1)(k − 2) · · · (k − n + 1).
• At the other extreme, the graph G = Kn with n vertices and no edges has chromatic polynomial k n ,
since every coloring is proper.
• If T is a tree with n vertices, then pick any vertex as the root; this imposes a partial order on the
vertices in which the root is 1̂ and each non-root vertex v is covered by exactly one other vertex p(v)
(its “parent”). There are k choices for the color of the root, and once we know f (p(v)) there are k − 1
choices for f (v). Therefore pT (k) = k(k − 1)n−1 . Qs
• If G has connected components G1 , . . . , Gs , then pG (k) = i=1 pGi (k). Equivalently, pG+H (k) =
pG (k)pH (k), where + denotes disjoint union of graphs.
Theorem 4.4.2. For every graph G
pG (k) = (−1)n−c k c · TG (1 − k, 0)
where n is the number of vertices of G and c is the number of components. In particular, pG (k) is a polynomial
function of k.
101
Proof. First, we show that the chromatic function satisfies the recurrence
pG (k) = k n if E = ∅; (4.7)
pG (k) = 0 if G has a loop; (4.8)
pG (k) = (k − 1)pG/e (k) if e is a coloop; (4.9)
pG (k) = pG\e (k) − pG/e (k) otherwise. (4.10)
We already know (4.7) and (4.8). Suppose e = xy is not a loop. Let f be a proper k-coloring of G \ e. If
f (x) = f (y), then we can identify x and y to obtain a proper k-coloring of G/e. If f (x) ̸= f (y), then f is a
proper k-coloring of G. So (4.10) follows.
This argument applies even if e is a coloop. In that case, however, the component H of G containing e
becomes two components H ′ and H ′′ of G \ e, whose colorings can be chosen independently of each other.
So the probability that f (x) = f (y) in any proper coloring is 1/k, implying (4.9).
The graph G \ e has n vertices and either c + 1 or c components, according as e is or is not a coloop.
Meanwhile, G/e has n − 1 vertices and c components. By induction,
n
k if E = ∅,
0 if e is a loop,
(−1)n−c k c TG (1 − k, 0) =
(1 − k)(−1)n+1−c k c TG/e (1 − k, 0) if e is a coloop,
(−1)n−c k c TG\e (1 − k, 0) + TG/e (1 − k, 0)
otherwise
n
k if E = ∅,
0 if e is a loop,
=
(k − 1)pG/e (k) if e is a coloop,
pG\e (k) − pG/e (k) otherwise
More generally, if G is a graph with n vertices and c components, then its graphic matroid M = M (G) has
rank n − c, whose associated geometric lattice is the connectivity lattice K(G). Combining Theorems 4.4.1
and 4.4.2 gives
pG (k) = k c χ(K(G); k).
An orientation is acyclic if it has no directed cycles. Let A(G) be the set of acyclic orientations of G, and let
a(G) = |A(G)|.
For example:
102
1. If G has a loop then a(G) = 0.
2. If G has no loops, then every total order on the vertices gives rise to an acyclic orientation: orient each
edge from smaller to larger vertex. Of course, different total orders can produce the same a.o.
3. If G has no edges than a(G) = 1. Otherwise, a(G) is even, since reversing all edges is a fixed-point
free involution on A(G).
4. Removing parallel copies of an edge does not change a(G), since all parallel copies would have to be
oriented in the same direction to avoid any 2-cycles.
5. If G is a forest then every orientation is acyclic, so a(G) = 2|E(G) .
6. If G = Kn then the acyclic orientations are in bijection with the total orderings, so a(G) = n!.
7. If G = Cn (the cycle of graph of length n) then it has 2n orientations, of which exactly two are not
acyclic, so a(Cn ) = 2n − 2.
Colorings and orientations are intimately connected. Given a proper coloring f : V (G) → [k], one can
naturally define an acyclic orientation by directing each edge from the smaller to the larger color. (So #2 in
the above list is a special case of this.) The connection between them is the prototypical example of what is
called combinatorial reciprocity.
A (compatible) k-pair for a graph G = (V, E) is a pair (O, f ), where O is an acyclic orientation of G and
f : V → [k] is a coloring such that f (x) ≤ f (y) for every directed edge x → y in D. Let C(G) = C(G, k) be
the set of compatible k-pairs of G (we can safely drop k from the notation)
Theorem 4.5.1 (Stanley’s Acyclic Orientation Theorem). For every graph G and positive integer k,
Proof. The second equality follows from Theorem 4.4.2, so we prove the first one. Let n = |G|.
If G has no edges then |C(G)| = k n = (−1)n (−k)n = (−1)n pG (−k), confirming (4.11).
If G has a loop then it has no acyclic orientations, hence no k-pairs for any k, so both sides of (4.11) are zero.
Let e = xy be an edge of G that is not a loop. Denote the left-hand side of (4.11) by p̄G (k). Then
Say that a pair (O, f ) ∈ C(G) is reversible (with respect to e) if reversing e produces a compatible pair
(O′ , f ); otherwise it is irreversible. (Reversibility is equivalent to saying that f (x) = f (y) and that G does
not contain a directed path from either endpoint of e to the other.) Let Crev (G) and Cirr (G) denote the sets
of reversible and irreversible compatible pairs, respectively.
If e is reversible, then contracting it to a vertex z and defining f (z)−f (x)−f (y) produces a compatible pair of
G/e. (The resulting orientation is acyclic because any directed cycle lifts to either a directed cycle in G, or an
oriented path between the endpoints of e, neither of which exists.) This defines a map ψ : Crev (G) → C(G/e),
which is 2-to-1 because ψ(O, f ) = ψ(O′ , f ). Moreover, ψ is onto: any (O, f ) ∈ C(G/e) can be lifted to
(Õf˜) ∈ C(G) by defining f˜(x) = f˜(y) = f (z) and orienting e in either direction (the acyclicity of O means
that there is no oriented path from either x or y to the other in Õ). We conclude that
|Crev (G)|
|C(G/e)| = . (4.12)
2
103
There is a map ω : C(G) → C(G − e) given by deleting e. I claim that ω is surjective, which is equivalent
to saying that it is always possible to extend any element of C(G − e) to C(G) by choosing an appropriate
orientation for e. (If f (x) < f (y), then O has no y, x-path by compatibility. If f (x) = f (y) and neither
orientation of e is acyclic, then O must contain a directed path from each of x, y to the other, hence is not
acyclic.) The map ω is 1-to-1 on Cirr (G) but 2-to-1 on Crev (G) (for the same reason as ψ). Therefore,
|Crev (G)|
|C(G − e)| = |Cirr (G)| + . (4.13)
2
as desired.
In particular, if k = 1 then there is only one choice for f and every acyclic orientation is compatible with it,
which produces the following striking corollary (often referred to as “Stanley’s theorem on acyclic orienta-
tions,” although Stanley himself prefers that name for the more general Theorem 4.5.1).
Theorem 4.5.2. The number of acyclic orientations of G is |pG (−1)| = TG (2, 0).
Combinatorial reciprocity can be viewed geometrically. For more detail, look ahead to Section 5.5 and/or
see a source such as Beck and Robins [BR15], but here is a brief taste.
Let G be a simple graph on n vertices. The graphic arrangement AG is the union of all hyperplanes in Rn
defined by the equations xi = xj where ij is an edge of G. The complement Rn \ AG consists of finitely
many disjoint open polyhedra (the “regions” of the arrangement), each of which is defined by a set of
inequalities, including either xi < xj or xi > xj for each edge. Thus each region naturally gives rise to an
orientation of G, and it is not hard to see that the regions are in fact in bijection with the acyclic orientations.
Meanwhile, a k-coloring of G can be regarded as an integer point in the cube [1, k]n ⊆ Rn , and a proper
coloring corresponds to a point that does not lie on any hyperplane in AG . In this setting, Stanley’s theorem
is an instance of something more general called Ehrhart reciprocity (which I will add notes on at some point).
Definition 4.6.1. A linear code C is a subspace of (Fq )n , where q is a prime power and Fq is the field of
order q. The number n is the length of C . The elements c = (c1 , . . . , cn ) ∈ C are called codewords. The
support of a codeword is supp(c) = {i ∈ [n] : ci ̸= 0}, and its weight is wt(c) = | supp(c)|. The weight
enumerator of C is the polynomial X
WC (t) = twt(c) .
c∈C
For example, let C be the subspace of F32 generated by the rows of the matrix
1 0 1
X= ∈ (F2 )3×2 .
0 1 1
104
The dual code C ⊥ is the orthogonal complement under the standard inner product. This inner product is
nondegenerate, i.e., dim C ⊥ = n − dim C . (Note, though, that a subspace and its orthogonal complement
can intersect nontrivially. A space can even be its own orthogonal complement, such as {00, 11} ⊆ F22 . This
does not happen over R, where the inner product is not only nondegenerate but also positive-definite, but
“positive” does not make sense over a finite field.) In this case, C ⊥ = {000, 111} and WC ⊥ (t) = 1 + t3 .
Theorem 4.6.2 (Curtis Greene, 1976). Let C be a linear code of length n and dimension r over Fq , and let M be the
matroid represented by the columns of a matrix X whose rows are a basis for C . Then
1 + (q − 1)t 1
WC (t) = tn−r (1 − t)r TM ,
1−t t
The proof is a deletion-contraction argument. As an example, if C = {000, 101, 011, 110} ⊆ F32 as above,
then the matroid
Mis U2 (3). Its Tutte polynomial is x2 + x + y, and Greene’s theorem gives WC (t) =
1+t 1
t(1 − t)2 TM 1−t , t = 1 + 3t2 as noted above (calculation omitted).
If X ⊥ is a matrix whose rows are a basis for the dual code, then the corresponding matroid M ⊥ is precisely
the dual matroid to M . We know that TM (x, y) = TM ⊥ (y, x) by (4.3), so setting s = (1 − t)/(1 + (q − 1)t) (so
t = (1 − s)/(1 + (q − 1)s); isn’t that convenient?) gives
1 + (q − 1)s 1
WC ⊥ (t) = tr (1 − t)n−r TM ,
1−s s
= tr (1 − t)n−r sr−n (1 − s)−r WC (s),
or rewriting in terms of t,
1 + (q − 1)tn 1−t
WC ⊥ (t) = WC
qr 1 + (q − 1)t
which is known as the MacWilliams identity and is important in coding theory.
4.7 Exercises
Problem 4.1. Give two proofs of the fact that TM ∗ (x, y) = TM (y, x), following the hints following equa-
tion (4.3).
Solution: Using deletion/contraction: We induct on |E|. The statement is clear when |E| = 0, for then TM =
TM ∗ = 1. If |E| > 0, let e ∈ E; then by Theorem 4.1.5,
xTM ∗ \e (y, x)
if e is a loop in M ∗
TM ∗ (y, x) = yTM ∗ /e (y, x) if e is a coloop in M ∗
if e is ordinary in M ∗
TM ∗ \e (y, x) + TM ∗ /e (y, x)
By induction and the fact that dualization interchanges loops and coloops, this becomes
xTM/e (x, y)
if e is a coloop in M
= yTM \e (x, y) if e is a loop in M
TM/e (x, y) + TM \e (x, y) if e is ordinary in M
105
which is precisely the recursive definition of TM (x, y).
Using the rank function: Let r and r∗ denote the rank functions in M and M ∗ respectively. Then
X ∗ ∗ ∗
TM ∗ (y, x) = (y − 1)r (E)−r (A) (x − 1)|A|−r (A)
A⊆E
X
= (y − 1)(r(E\E)+|E|−r(E))−(r(E\A)+|A|−r(E)) (x − 1)|A|−(r(E\A)+|A|−r(E)) (by Problem 3.8)
A⊆E
X
= (y − 1)|E|−r(E\A)−|A| (x − 1)−r(E\A)+r(E)
A⊆E
X
= (y − 1)|Z|−r(Z) (x − 1)r(E)−r(Z) (setting Z = E \ A)
Z⊆E
= TM (x, y).
Problem 4.2. An orientation of a graph is called totally cyclic if every edge belongs to a directed cycle.
Prove that the number of totally cyclic orientations of G is TG (0, 2).
Solution: Let Θ(G) denote the set of strong orientations of G, and let θ(G) = |Θ(G)|. The goal is to use the
Tutte Recipe Theorem for Graphs (Theorem 4.2.2).
Note first that θ(G) = 1 if G has no edges; θ(G) = 0 if G has a coloop; and θ(G − e) = 2θ(G) if e is a loop
(since even a loop edge has two possible orientations). The interesting case is when e is an ordinary edge.
Suppose that O is a strong orientation of G and that e is an ordinary edge, oriented as ⃗e = xy ⃗ in O. Call
⃗ gives a strong orientation O′ ; otherwise, O is e-irreversible. Observe
O e-reversible if replacing ⃗e with yx
that
Contracting any edge in a strong orientation produces a strong orientation (since contracting an edge in an
oriented cycle produces a cycle), so we have a map Θ(G) → Θ(G/e). The fibers of this map are the e-pairs
mentioned previously, together with the e-irreversible orientations (considered as singleton sets). On the
other hand, I claim that the map is surjective. Given a strong orientation of G/e, we can pull it back to an
orientation of G − e, in which oriented cycles pull back either to cycles, or to directed paths between x to y.
If no pullbacks are paths then e can be oriented in other way; if at least one pullback is a path from x to y,
then e can be oriented as yx,
⃗ and reversely.
Therefore,
Now applying the Tutte Recipe Theorem for Graphs with a = 1, X = 0, Y = 2, b = 1, c = 1 gives the
desired result.
106
Problem 4.3. Let G be a finite graph with n vertices, r edges, and k components. Fix an orientation O
on E(G). Let I(v) (resp., O(v)) denote the set of edges entering (resp., leaving) each vertex v. Let q be a
positive integer and Zq = Z/qZ. A nowhere-zero q-flow (or q-NZF) on G (with respect to O) is a function
ϕ : E(G) → Zq \ {0} satisfying the conservation law
X X
ϕ(e) = ϕ(e)
e∈I(v) e∈O(v)
for every v ∈ V (G). Let FGO (q) denote the set of nowhere-zero q-flows and fG
O
(q) = |FGO (q)|.
O
(i) Prove that fG (q) depends only on the graph G, not on the choice of orientation (so we are justified in
writing fG (q)).
(Interestingly, Z/qZ can be replaced with any abelian group of cardinality q without affecting the result.)
Solution: (i) Fix a graph G and orientation O and let ϕ ∈ FGO (q). Fix e ∈ E and define ϕ′ by ϕ′ (e) = −ϕ(e)
′
and ϕ′ (f ) = ϕ(f ) for f ̸= e. Then ϕ′ ∈ FGO (q), where O′ is obtained by reversing e. (The conservation
law is unchanged for most vertices v; for the endpoints of v, the summand ϕ(e) on one side is replaced by
′
−ϕ(e) on the other.) This operation is an involution, hence gives a bijection between FGO (q) and FGO (q).
Any orientation can be obtained from any other by repeatedly reversing edges, so we conclude that fG (q)
is well-defined.
(ii) We are going to apply the Recipe Theorem. By fiat, fG (q) = 1 if E(G) = ∅, since there is only one
function with domain the empty set.
If e is a loop, say at vertex x, then the conservation equations are unchanged by deleting it (since ϕ(e)
appears on both sides for v = x and on neither side for v ̸= x). So every q-NZF ϕ on G consists of a q-NZF
on G − e together with a choice for ϕ(e), which can be any of the q − 1 nonzero values of Z/qZ.
If G has a coloop (cut-edge) e, then G has no nonzero q-flows, for the following reason. Say e is oriented as
x → y. Let G′ be the component of G containing e and let H be the component of G′ − e containing y. If ϕ
is a q-NZF, then
X X X X X
0= ϕ(a) − ϕ(a) =
ϕ(a)−
ϕ(a) = ϕ(e) ̸= 0
v∈V (H) a∈I(v) a∈O(v) a∈E(G): a∈E(G):
a enters some v∈V (H) a leaves some v∈V (H)
Finally, suppose e is an ordinary edge. Let ϕ ∈ FG/e (q), let x, y be the endpoints of e in G (say e = x → y),
and let z be the vertex of G/e obtained by identifying x, y. If we lift ϕ to a flow ϕ̃ on G (leaving ϕ̃(e)
undefined for the moment), then conservation certainly holds for all v ̸∈ {x, y}. For x and y, observe that
107
Call this number c. We now see if that c ̸= 0, then we can define ϕ̃(e) = ϕ̃(x → y) = c to obtain a q-NZF on
G. On the other hand, if c = 0, then ϕ̃ is just a q-NZF on G − e. So we have constructed a map
that is certainly 1-1. On the other hand, any q-NZF ψ on either G or G\e can be made into a q-NZF on G/e
in the obvious way (forgetting the value of ψ(e) in the former case). If ψ ∈ FG\e (q) then conservation in
G\e immediately implies it for G/e, while if ψ ∈ FG (q) then equating the conservation laws at x and y and
cancelling ψ(e) produces (4.14). So the map in (4.15) is a bijection and
The result now follows from applying the Recipe Theorem with X = 0, Y = q − 1, a = 1, b = −1.
Problem 4.4. Let G = (V, E) be a graph with n vertices and c components. For a vertex coloring f : V → P,
let i(f ) denote the number of “improper” edges, i.e., whose endpoints are assigned the same color. The
(Crapo) coboundary polynomial of G is
X
χ¯ G (q; t) = q −c ti(f ) .
f :V →[q]
This is evidently a stronger invariant than the chromatic polynomial of G, which can be obtained as q χ¯ G (q, 0).
In fact, the coboundary polynomial provides the same information as the Tutte polynomial. Prove that
−
n−c q + t 1
χ¯ G (q; t) = (t − 1) TG , t
t−1
Solution: The goal is to use the Tutte Recipe Theorem for Graphs (Theorem 4.2.2).
1. If G has no edges, then χ¯ G (q; t) = q −c q n = 1, since all colorings are proper and c = n.
2. If e is a loop, then it is an improper edge for every coloring, so χ¯ G (q; t) = tχ¯ G−e (q; t).
3. If e = xy is an ordinary edge, then c does not change under deletion or contraction. Since the quantity i(f )
depends on what graph is being referred to, we will write i, i′ , i′′ for G, G − e, G/e respectively. Therefore,
X ′
χ¯ G (q, t) − χ¯ G−e (q, t) = q −c (ti(f ) − ti (f ) ) (4.16)
f :V (G)→[q]
X ′
= q −c (ti(f ) − ti (f ) )
f :V →[q]
f (x)=f (y)
108
(where V /xy = V (G/e) is V with x, y e identified), and therefore
4. If e = xy is a cut-edge, then we can proceed similarly, except that G − e has one fewer component than
G, so that the left side of (4.16) must be χ¯ G (q, t) − q χ¯ G−e (q, t), and we obtain
In order to use the recipe theorem, we really need an expression in terms of χ¯ G−e . Let H be the component
containing one of the endpoints of e. Consider the equivalence relation on colorings given as follows: f ∼ f ′
if f ′ can be obtained from f by choosing some constant a ∈ Z/q and setting f ′ (v) = f (v) + a for v ∈ H,
and leaving f unchanged outside H. Note that each equivalence class has size q and that both the number
of improper edges, and the color assigned to y, is constant on each equivalence class. Moreover, precisely
one coloring in each equivalence class has f (x) = f (y) and therefore corresponds to a coloring of G/e. On
the other hand, G − e has one more component than G/e, so in fact χ¯ G−e (q, t) = χ¯ G/e (q, t), and combined
with (4.17) we get
χ¯ G (q, t) = (q + t − 1)χ¯ G−e (q, t).
To summarize:
1 if E = ∅,
(q + t − 1)χ¯
if e ∈ E is a coloop,
G\e
χ¯ G =
t ¯
χ
G\e
if e ∈ E is a loop,
¯
χ G\e + (t − 1)χ¯ G/e if e ∈ E is ordinary,
Therefore, applying the Tutte Recipe Theorem for Graphs with a = 1, X = q + t − 1, Y = t, b = 1, c = t − 1
gives the desired formula.
Problem 4.5. Let M be a matroid on E and let 0 ≤ p ≤ 1. The reliability polynomial RM (p) is the
probability that the rank of M stays the same when each ground set element is independently retained with
probability p and deleted with probability 1−p. (In other words, we have a family of i.i.d. random variables
{Xe : e ∈ E}, each of which is 1 with probability p and 0 with probability 1 − p. Let A = {e ∈ E : Xe = 1}.
Then RM (p) is the probability that r(A) = r(E).) Give a formula for RM (p) in terms of the Tutte polynomial,
using
(a) the definition of the Tutte polynomial as the corank/nullity generating function;
(b) the Tutte Recipe Theorem.
Recall that X
TM (x, y) = (x − 1)r(E)−r(A) (y − 1)|A|−r(A) .
A⊆E
109
In order to make sure that we are only including sets of full rank, we set x = 1 to get
X
TM (1, y) = (1 − 1)r(E)−r(A) (y − 1)|A|−r(A)
A⊆E
X
= (y − 1)|A|−r(E)
A⊆E
r(A)=r(E)
X
= (y − 1)−r(E) (y − 1)|A| (4.19)
A⊆E
r(A)=r(E)
p p 1
then comparing (4.19) with (4.20) it becomes evident that we should set y − 1 = 1−p , i.e., y = 1 + 1−p = 1−p .
Plugging this value of y into (4.19), we get
−r(E) |A|
p X p
TM (1, 1/(1 − p)) =
1−p 1−p
A⊆E
r(A)=r(E)
−r(E)
p
= (1 − p)−|E| RM (p)
1−p
and so
|E|−r(E) r(E) 1
RM (p) = (1 − p) p TM 1, .
1−p
(b) We claim that the reliability polynomial satisfies the following recurrence for all e ∈ E:
1 if E = ∅,
R
M \e (p) if e is a loop,
RM (p) =
pRM/e (p) if e is a coloop,
pRM/e (p) + (1 − p)RM \e (p) if e is ordinary,
The first case is trivial. For the second and third, we do not care whether or not loops are retained, but any
coloops must be retained. The last case follows from the observation that
RM (p) = Pr[e is retained and A \ e spans M/e] + Pr[e is dropped and A \ e spans M ].
Given the recurrence, the result follows from the Recipe Theorem.
Problem 4.6. Prove Merino’s theorem on critical configurations of the chip-firing game. (This needs de-
tails!)
Solution: To be written
Problem 4.7. Prove Theorem 4.3.4.
110
Problem 4.8. Prove Theorem 4.6.2.
Solution: To be written
Much, much more about the Tutte polynomial can be found in [BO92], the MR review of which begins,
“The reviewer, having once worked on that polynomial himself, is awed by this exposition of its present
importance in combinatorial theory.” (The reviewer was one W.T. Tutte.)
111
Chapter 5
Hyperplane Arrangements
An excellent source for the combinatorial theory of hyperplane arrangements is Stanley’s book chapter
[Sta07], which is accessible to newcomers, and includes a self-contained treatment of topics such as the
Möbius function and characteristic polynomial. Another canonical (but harder) source is the monograph
by Orlik and Terao [OT92].
Definition 5.1.1. Let k be a field, typically either R or C, and let n ≥ 1. A linear hyperplane in kn is a
vector subspace of codimension 1. An affine hyperplane is a translate of a linear hyperplane. A hyperplane
arrangement A ⊆ kn is a finite set of (distinct) hyperplanes H1 , . . . , Hk ⊆ kn . The number n is called the
dimension of A, and the space kn is its ambient space. The intersection poset L(A) is the poset of all
nonempty intersections
T of subsets of A, ordered by reverse inclusion. If B ⊆ A is a subset of hyperplanes,
we write ∩B for H∈B H. The characteristic polynomial of A is
X
χA (t) = µ(0̂, x)tdim x . (5.1)
x∈L(A)
This is essentially the same as the characteristic polynomial of the poset L(A), up to a correction factor that
we will explain soon.
Example 5.1.2. Two line arrangements in R2 are shown in Figure 5.1. The arrangement A1 consists of the
lines x = 0, y = 0, and x = y. The arrangement A2 consists of the four lines ℓ1 , ℓ2 , ℓ3 , ℓ4 given by the
equations y = 1, x = y, x = −y, y = −1 respectively. The intersection posets L(A1 ) and L(A2 ) are shown in
Figure 5.2; the characteristic polynomials are t2 − 3t + 2 and t2 − 4t + 5 respectively. ◀
Example 5.1.3. The Boolean arrangement Booln (or coordinate arrangement) consists of the n coordinate
hyperplanes in n-space. Its intersection poset is the Boolean lattice Booln (I make no apologies for abus-
ing notation by referring to the arrangement and the poset with the same symbol). More generally, any
arrangement whose intersection poset is Boolean might be referred to as a Boolean arrangement. ◀
n
Example 5.1.4. The braid arrangement Brn consists of the 2 hyperplanes xi = xj in n-space. Its intersec-
tion poset is naturally identified with the partition lattice Πn . This is simply because any set of equalities
among x1 , . . . , xn defines an equivalence relation on [n], and certainly every equivalence relation can be
obtained in this way. For instance, the intersection poset of Br3 is as follows:
112
`1
`2
`3
A1 `4 A2
{0} • • • • •
L(A1 ) R2 L(A2 ) R2
x=y=z 123
R3 1|2|3
Note that the poset Πn = L(Br3 ) has characteristic polynomial t2 − 3t + 2, but the arrangement Br3 has
characteristic polynomial t3 − 3t2 + 2t. ◀
Figure 5.3 shows some hyperplane arrangements in R3 . Note that every hyperplane in Brn contains the line
x1 = x2 = · · · = xn ,
so projecting R4 along that line allows us to picture Br4 as an arrangement ess(Br4 ) in R3 . (The symbol
“ess” means essentialization, to be defined precisely soon.) The second two figures were produced using the
computer algebra system Sage [S+ 14].
The poset L(A) is the fundamental combinatorial invariant of A. Some easy observations:
1. If T : Rn → Rn is an invertible linear transformation, then L(T (A)) ∼= L(A), where T (A) = {T (H) : H ∈
A}. In fact, the intersection poset is invariant under any affine transformation. (The group of affine trans-
formations is generated by the invertible linear transformations together with translations.)
2. The poset L(A) is a meet-semilattice, with meet given by ∩B ∧ ∩C = ∩(B ∩ C) for all B, C ⊆ A. Its 0̂
element is ∩∅, which by convention is kn .
113
Bool3 Br3 ess(Br4)
Figure 5.3: Three hyperplane arrangements in R3 .
3. L(A) is ranked, with rank function r(X) = n − dim X. To see this, observe that each covering relation
X ⋖ Y comes from intersecting an affine linear subspace X with a hyperplane H that neither contains nor
is disjoint from X, so that dim(X ∩ H) = dim X − 1.
4. L(A) has a 1̂ element if and only if the center ∩A is nonempty. Such an arrangement is called central. In
this case L(A) is a lattice (and may be referred to as the intersection lattice of A). Since translation does not
affect whether an arrangement is central (or indeed any of its combinatorial structure), we will typically
assume that ∩A contains the zero vector, which is to say that every hyperplane in A is a linear hyperplane
in kn . (So an arrangement is central if and only if it is a translation of an arrangement of linear hyperplanes.)
5. When A is central, the lattice L(A) is geometric. It is atomic by definition, and it is submodular because
it is a sublattice of the chain-finite modular lattice L(kn )∗ (the lattice of all subspaces of kn ordered by
reverse inclusion). The associated matroid M (A) = M (L(A)) is represented over k by any family of vectors
{nH : H ∈ A} where nH is normal to H. (That is, H ⊥ = k⟨nH ⟩ with respect to some fixed non-degenerate
bilinear form on kn .) Any normals will do, since the matroid is unchanged by scaling the nH independently.
Therefore, all of the tools we have developed for looking at posets, lattices and matroids can be applied to
study hyperplane arrangements.
The dimension of an arrangement is not a combinatorial invariant; that is, it cannot be extracted from
the intersection poset. If Br4 were a “genuine” 4-dimensional arrangement then we would not be able to
represent it in R3 . However, we can do so because the center of Br4 has positive dimension, so squashing the
center to a point reduces the ambient dimension without changing the intersection poset. This observation
motivates the following definition.
Definition 5.1.5. Let A ⊆ kn be an arrangement and let N (A) = k⟨nH : H ∈ A⟩, where nH is normal to H.
The essentialization of A is the arrangement
This is why there is an extra power of k in the characteristic polynomial of the arrangement (as opposed to
114
the intersection poset), so that it can record the dimension of A. Specifically,
χA (t) = tdim N (A) χL(A) (t) = tdim A−rank A χL(A) (t). (5.2)
The two polynomials coincide for essential arrangements. For example, rank Brn = dim ess(Brn ) = n − 1,
and rank AG = r(G) = |V (G)| − c, where c is the number of connected components of G.
If A is linear, then we could define the essentialization by setting V = N (A)⊥ = ∩A, and then defining
ess(A) = {H/V : H ∈ A} ⊆ kn /V . Thus A is essential if and only if ∩A = 0. Moreover, if A is linear then
rank(A) is the rank of its intersection lattice — so rank is a combinatorial invariant, unlike dimension.
Example 5.1.6. If G = (V, E) is a simple graph on vertex set V = [n], then the corresponding graphic
arrangement AG is the subarrangement of Brn consisting of those hyperplanes xi = xj for which ij ∈ E.
Thus Brn itself is the graphic arrangement of the complete graph Kn .
Moreover, the intersection poset of AG is precisely the connectivity lattice K(G) defined in Example 1.2.3.
Precisely, each subset B ⊆ AG corresponds to a set of edges EB that induce some partition π ∈ K(G), and
the corresponding intersection is
In particular dim ∩B = |π|, so the definition (5.1) of the characteristic polynomial of AG becomes
X
χA (t) = µ(0̂, π)t|π| = pG (t),
π∈K(G)
the chromatic polynomial of G (by (2.5)). In §5.4, we will see a more explicit way in which the characteristic
polynomial enumerates colorings. ◀
There are two natural operations that go back and forth between central and non-central arrangements,
called projectivization and coning.
Let k be a field and n ≥ 1. The set of lines through the origin in kn is called n-dimensional projective space
over k and denoted by Pn−1 k. For example, if k = R, we can regard Pn−1 R as the unit sphere Sn−1 with
opposite points identified. (In particular, it is an (n − 1)-dimensional manifold, although it is orientable
only if n is even.)
Algebraically, write x ∼ y if x and y are nonzero scalar multiples of each other. Then ∼ is an equivalence
relation on kn \{0}, and Pn−1 is the set of equivalence classes. In particular, each linear hyperplanes H ⊂ kn
correspond to a set of equivalence classes that form an affine hyperplane proj(H) ⊆ Pn−1 k.
Definition 5.1.7. Let A ⊆ kn be a central arrangement. Its projectivization proj(A) is the affine arrange-
ment {proj(H) | H ∈ A} in Pn−1 k.
Projectivization supplies a nice way to draw central 3-dimensional real arrangements. Let S be the unit
sphere, so that H ∩ S is a great circle for every H ∈ A; then regard H0 ∩ S as the equator and project the
northern hemisphere into your piece of paper. Several examples as shown below. Of course, a diagram of
proj(A) only shows the upper half of A; we can recover A from proj(A) by “reflecting the interior of the
disc to the exterior” (Stanley); e.g., for the Boolean arrangement A = Bool3 , the picture is as shown in the
fourth figure below. In general, r(proj(A)) = 21 r(A).
115
proj(Bool3 ) proj(Br3 ) proj(ess(Br4 )) Bool3 from proj(Bool3 )
The operation of coning is a sort of inverse of projectivization. It lets us turn a non-central arrangement into
a central arrangement, at the price of increasing the dimension by 1.
Definition 5.1.8. Let A ⊆ kn be a hyperplane arrangement, not necessarily central. The cone cA is the
central arrangement in kn+1 defined as follows:
• Geometrically: Make a copy of A in kn+1 , choose a point p not in any hyperplane of A, and replace
each H ∈ A with the affine span H ′ of p and H (which will be a hyperplane in kn+1 ). Then, toss in
one more hyperplane containing p and in general position with respect to every H ′ .
• Algebraically: For H = {x : L(x) = ai } ∈ A (with L a homogeneous linear form on kn and ai ∈ k),
construct a hyperplane H ′ = {(x1 , . . . , xn , y) : L(x) = ai y} ⊆ kn+1 in cA. Then, toss in the hyperplane
y = 0.
For example, if A consists of the points x = 0, x = −3 and x = 1 in R1 (shown in red), then cA consists of
the lines x = y, x = −5y, x = 3y, and y = 0 in R2 (shown in blue).
y=1
y=0
Let A ⊆ Rn be a real hyperplane arrangement. The regions of A are the connected components of Rn \ A.
Each component is the interior of a (bounded or unbounded) polyhedron; in particular, it is homeomorphic
to Rn . We call a region relatively bounded if the corresponding region in ess(A) is bounded. (If A is not
essential then every region is unbounded, because it contains a translate of W ⊥ , where W is the space
defined in Definition 5.1.5. Therefore passing to the essentialization is necessary to make the problem of
counting bounded regions nontrivial for all arrangements.) Let
116
Example 5.2.1. For the arrangements A1 and A2 shown in Example 5.1.2,
r(A1 ) = 6, r(A2 ) = 10,
b(A1 ) = 0, b(A2 ) = 2. ◀
Example 5.2.2. The Boolean arrangement Booln consists of the n coordinate hyperplanes in Rn . It is
a central, essential arrangement whose intersection lattice is the Boolean lattice of rank n; accordingly,
χBooln (t) = (t − 1)n . The complement Rn \ Booln is {(x1 , . . . , xn ) : xi ̸= 0 for all i}, and the connected com-
ponents are the open orthants, specified by the signs of the n coordinates. Therefore, r(Booln ) = 2n and
b(Booln ) = 0. ◀
Example 5.2.3. Let A consist of m lines in R2 in general position: that is, no two lines are parallel and no
three are coincident. Draw the dual graph G, whose vertices are the regions of A, with an edge between
every two regions that share a common border.
Let r = r(A) and b = b(A), and let v, e, f denote the numbers of vertices, edges and faces of G, respectively.
(In the example above, (v, e, f ) = (11, 16, 7).) Each bounded face of G is a quadrilateral that contains exactly
one point where two lines of A meet, and the unbounded face is a cycle of length r − b. Therefore,
v = r, (5.3a)
m2 − m + 2
m
f = 1+ = (5.3b)
2 2
4(f − 1) + (r − b) = 2e. (5.3c)
Moreover, the number r − b of unbounded regions of A is just 2m. (Take a walk around a very large circle.
You will enter each unbounded region once, and will cross each line twice.) Therefore, from (5.3c) and
(5.3b) we obtain
e = m + 2(f − 1) = m2 . (5.3d)
Euler’s formula for planar graphs says that v −e+f = 2. Substituting in (5.3a), (5.3b) and (5.3d) and solving
for r gives
m2 + m + 2
r =
2
and therefore
m2 − 3m + 2 m−1
b = r − 2m = = .
2 2
◀
117
Example 5.2.4. The braid arrangement Brn consists of the n2 hyperplanes Hij = {x : xi = xj } in Rn .
The complement Rn \ Brn consists of all vectors in Rn with no two coordinates equal, and the connected
components of this set are specified by the ordering of the set of coordinates as real numbers:
y=x
y<x<z x<y<z
z=y
y<z<x x<z<y
z=x
z<y<x z<x<y
Therefore, r(Brn ) = n!. (Stanley: “Rarely is it so easy to compute the number of regions!”) Furthermore,
Note that the braid arrangement is central but not essential; its center is the line x1 = x2 = · · · = xn , so its
rank is n − 1. ◀
Example 5.2.5. Let G = (V, E) be a simple graph with V = [n], and let AG be its graphic arrangement
(see Example 5.1.6). The characteristic polynomial of L(AG ) is precisely the chromatic polynomial of G (see
Section 4.4). We will see another explanation for this fact later; see Example 5.4.4.
The regions of Rn \ AG are the open polyhedra whose defining inequalities include either xi < xj or xi > xj
for each edge ij ∈ E. Those inequalities give rise to an orientation of G, and it is not hard to check that this
correspondence is a bijection between regions and acyclic orientations. Hence
Example 5.2.5 motivates the main result of this section, historically the first major theorem about hyperplane
arrangements.
Theorem 5.3.1 (Zaslavsky’s Theorem [Zas75]). Let A be a real hyperplane arrangement, and let χA be the char-
acteristic polynomial of its intersection poset. Then
118
1. Show that r and b satisfy restriction/contraction recurrences in terms of associated hyperplane ar-
rangements A′ and A′′ (Prop. 5.3.3).
2. Rewrite the characteristic polynomial χA (k) as a sum over central subarrangements of A (the “Whit-
ney formula”, Prop. 5.3.4).
3. Show that the Whitney formula obeys a restriction/contraction recurrence (Prop. 5.3.5) and compare
it with those for r and b.
Let x ∈ L(A), i.e., x is a nonempty affine space formed by intersecting some of the hyperplanes in A. Define
Ax = {H ∈ A : H ⊇ x},
(5.6)
Ax = {H ∩ x : H ∈ A \ Ax }.
In other words, Ax is obtained by deleting the hyperplanes not containing x, while Ax is obtained by re-
stricting A to x so as to get an arrangement whose ambient space is x itself. The notation is mnemonic:
L(Ax ) and L(Ax ) are isomorphic respectively to the principal order ideal and principal order filter gener-
ated by x in L(A). That is,
L(Ax ) ∼
= {y ∈ L(A) : y ≤ x}, L(Ax ) ∼
= {y ∈ L(A) : y ≥ x}.
Example 5.3.2. Let A be the 2-dimensional arrangement shown on the left, with the line H and point p as
shown. Then Ap and AH are shown on the right.
H H
q r r H
U
p p p
L
s t s
K K
A Ap AH
The lattice L(A) and its subposets (in this case, sublattices) L(Ap ) and L(AH ) are shown below.
s r p t q s r p t q
U H K L U H K L
L(Ap ) L(AH )
Let M (A) be the matroid represented by normal vectors {nH : H ∈ A}. Fix a hyperplane H ∈ A and let
A′ = A \ H, A′′ = AH . (5.7)
119
Proposition 5.3.3. The invariants r and b satisfy the following recurrences:
′ ′′
1. r(A) = r(A
) + r(A ).
0 if rank A = rank A′ + 1 (i.e., if nH is a coloop in M (A)),
2. b(A) =
b(A′ ) + b(A′′ ) if rank A = rank A′ (i.e., if it isn’t).
Proof. (1) Consider what happens when we add H to A′ to obtain A. Some regions of A′ will remain the
same, while others will be split into two regions.
unsplit
split split
unsplit
0 A
A
unsplit unsplit
split
unsplit
H
Let S and U be the numbers of split and unsplit regions of A′ (in the figure above, S = 2 and U = 4).
The unsplit regions each contribute 1 to r(A). The split regions each contribute 2 to r(A), but they also
correspond bijectively to the regions of A′′ . (See, e.g., Example 5.3.2.) So
and so r(A) = r(A′ ) + r(A′′ ), proving the first assertion of Proposition 5.3.3. By the way, if (and only if) H
is a coloop then it borders every region of A, so r(A) = 2r(A′ ) in this case.
If rank A = rank A′ + 1, then N (A′ ) ⊊ Rn , i.e., A′ is not essential. In that case, every region of A′ must
contain a translate of the nonempty vector space N (A′ )⊥ , which gets squashed down to a point upon
essentialization. In particular, every region of A contains half of that translate, hence is unbounded, so
b(A) = 0.
If rank A = rank A′ , then the relatively bounded regions of A come in a few different flavors.
120
W W
X
s Z
H XY
Y ZU
U
A A0
Contributions to. . .
Description b(A) b(A′ ) b(A′′ )
(W) bounded regions that don’t touch H 1 1 0
(X, Y) pairs of bounded regions separated by H 2 1 1
(Z) bounded, neighbor across H is unbounded 1 0 1
In all cases the contribution to b(A) equals the sum of those to b(A′ ) and b(A′′ ), establishing the second
desired recurrence.
Proposition 5.3.3 looks a lot like a Tutte polynomial deletion/contraction recurrence. This suggests that we
should be able to extract r(A) and b(A) from the characteristic polynomial χA . The first step is to find a
more convenient form for the characteristic polynomial.
Proposition 5.3.4 (Whitney formula for χA ). For any hyperplane arrangement A,
X
χA (t) = (−1)|B| tdim A−rank B .
central B⊆A
Whitney [Whi32b, Whi32a] gave an equivalent formula for the chromatic polynomial of a graph, well be-
fore anyone was talking about crosscuts or hyperplane arrangements in those terms. So the more general
version gets his name on it as well. I don’t know whether the modern proof given below specializes to
Whitney’s proof.
Proof. For x ∈ L(A), consider the interval [0̂, x] as a sublattice of L(A). Its atoms are the hyperplanes of A
containing x, and they form a lower crosscut of [0̂, x]. Therefore,
X
χA (t) = µ(0̂, x)tdim x
x∈L(A)
X X
= (−1)|B| tdim x
T
x∈L(A) B⊆A: x= B
121
(by the second form of Rota’s crosscut theorem (Thm. 2.4.9); note that 1̂[0̂,x] = x)
X T
= (−1)|B| tdim( B)
T
B⊆A: B̸=0
X
= (−1)|B| tdim A−rank B
central B⊆A
as desired. Note that the empty subarrangement is considered central for the purpose of this formula (since
by convention its intersection is Rdim A ), corresponding to the summand x = 0̂ and giving rise to the leading
term tdim A of χA (t).
Proposition 5.3.5. Let A be a hyperplane arrangement in kn . As before, let H ∈ A and define A′ , A′′ as in (5.7).
Then χA (t) = χA′ (t) − χA′′ (t).
Then SU M 1 = χA′ (t) (it is just Whitney’s formula for A′ ), so it remains to show that SU M 2 = −χA′′ (t).
This is a little trickier, because different hyperplanes in A can have the same intersection with H, which
means that multiple subarrangements of A can give rise to the same subarrangement of A′′ .
Label the hyperplanes of A′′ (which, remember, are codimension-1 subspaces of H) as K1 , . . . , Ks . Then
A′′ is the union of the pairwise-disjoint sets Ai = {J ∈ A : J ∩ H = Ki }, for i ∈ [s]. Each arrangement B
arising as a summand of SU M 2 gives rise to a central subarrangement of A′′ , namely
π(B) = {J ∩ H : J ∈ B},
122
(to see this, expand the product and observe that equals the
P inner sum in the previous line; the outer minus
sign is contributed by H, which is an element of B). But ∅̸=Bi ⊆Ai (−1)|Bi | = −1, because it is the binomial
expansion of (1 − 1)|Ai | = 0, with one +1 term (namely Bi = ∅) removed. (Note that Ai ̸= ∅.) Therefore, the
whole thing boils down to X ′′ ′′
− (−1)|B | tdim H−rank B
B′′
We can now finish the proof of the main result. We have already done the hard work, and just need to put
all the pieces together.
Proof of Zaslavsky’s Theorem 5.3.1. Let r̃(A) and b̃(A) denote the numbers on the right-hand sides of (5.4)
and (5.4).
If |A| = 1, then L(A) is the lattice with two elements, namely Rn and a single hyperplane H, and its
characteristic polynomial is tn − tn−1 . Thus r̃(A) = (−1)n ((−1)n − (−1)n−1 ) = 2 and b̃(A) = −(1 − 1) = 0,
which match r(A) and b(A).
For the general case, we just need to show that r̃ and b̃ satisfy the same recurrences as r and b (see
Prop. 5.3.3). First,
As for b̃, if rank A = rank A′ + 1, then in fact A′ and A′′ have the same essentialization, hence the same
rank, and their characteristic polynomials only differ by a factor of t. The deletion/restriction recurrence
(Prop. 5.3.5) therefore implies b̃(A) = 0.
On the other hand, if rank A = rank A′ , then rank A′′ = rank A − 1 and a calculation similar to that for r̃
(replacing dimension with rank) shows that b̃(A) = b̃(A′ ) + b̃(A′′ ).
Corollary 5.3.7. Let A ⊆ Rn be a central hyperplane arrangement and let M = M (A) be the matroid represented
by normals. Then r(A) = TM (2, 0) and b(A) = 0.
Proof. Combine Zaslavsky’s theorem with the formula χA (t) = (−1)n TM (1 − t, 0) which needs to be
proved!, and use the fact that TM (0, 0) = 0 for any matroid M with nonempty ground set.
Remark 5.3.8. The formula for r(A) could be obtained from the Tutte Recipe Theorem (Thm. 4.2.1). But
this would not work for b(A), which is not an invariant of M (A). (The matroid M (A) is not as meaningful
when A is not central, which is precisely the case that b(A) is interesting.)
123
Example 5.3.9. Let s ≥ n, and let A be an arrangement of s linear hyperplanes in general position in Rn ;
that is, every k hyperplanes intersect in a space of dimension n − k (or 0 if k > n). Equivalently, the
corresponding matroid M is Un (s), whose rank function r : 2[s] → N is given by r(A) = min(n, |A|).
Therefore,
X
r(A) = TM (2, 0) = (2 − 1)n−r(A) (0 − 1)|A|−r(A)
A⊆[s]
X
= (−1)|A|−r(A)
A⊆[s]
s
X s
= (−1)k−min(n,k)
k
k=0
n s
X s X s
= + (−1)k−n
k k
k=0 k=n+1
n
X s s
k−n
X s
= (1 − (−1) ) + (−1)k−n
k k
k=0 k=0
| {z }
=0
s s s
= 2 + + + ··· .
n−1 n−3 n−5
Notice that this is not the same as the number of regions formed by s affine lines in general position in R2 .
The calculation of r(A) and b(A) for that arrangement is left to the reader (Problem 5.1).
Corollary 5.3.10. Let A be an arrangement in which no two hyperplanes are parallel. Then A has at least one
relatively bounded region if and only if it is noncentral. Prove this and find a place for it — assuming it is true.
The non-parallel assumption is necessary since the conclusion fails for the arrangement with hyperplanes
x = 0, y = 0, y = 1.
The following very important result is implicit in the work of Crapo and Rota [CR70] and was stated ex-
plicitly by Athanasiadis [Ath96]:
Theorem 5.4.1. Let Fq be the finite field of order q, and let A ⊆ Fnq be a hyperplane arrangement. Then
|Fnq \ A| = χA (q).
This result gives a combinatorial interpretation of the values of the characteristic polynomial. In practice,
it is often used to calculate the characteristic polynomial of a hyperplane arrangement by counting points
in its complement over Fq (which can be regarded as regions of the complement, if you endow Fnq with the
discrete topology).
124
Proof #1. By inclusion-exclusion,
X \
|Fnq \ A| = (−1)|B| B .
B⊆A
X
|Fnq \ A| = (−1)|B| q n−rank B
central B⊆A
Proof #2. Start with the definition of the characteristic polynomial, letting r be the rank function in L(A):
X
χA (q) = µ(0̂, x)q n−r(x)
x∈L(A)
X
= µ(0̂, x)q dim x
x∈L(A)
X
= µ(0̂, x)|x|
x∈L(A)
X X
= µ(0̂, x)
p∈Fn
q x∈L(A): p∈x
X X
= µ(0̂, x)
p∈Fn
q x∈[0̂,yp ]
T
where yp = H⊇p H. By definition of the Möbius function, the parenthesized sum is 1 if yp = 0̂ and 0
otherwise. Therefore
This fact has a much more general application, which was systematically mined by Athanasiadis, e.g.,
[Ath96].
Definition 5.4.2. Let A ⊆ Rn be an integral hyperplane arrangement (i.e., whose hyperplanes are defined
by equations with integer coefficients). For a prime p, let Ap = A ⊗ Fp be the arrangement in Fnp defined by
regarding the equations in A as equations over Fp . We say that A reduces correctly modulo p if L(Ap ) ∼ =
L(A). (We need only consider the prime case, since if q is a power of p, then L(Aq ) = L(Ap ).)
A sufficient condition for correct reduction is that no minor of the matrix of normal vectors is a nonzero
multiple of p (so that rank calculations are the same over Fp as over Z). In particular, if we choose p larger
than the absolute value of any minor of M , then each set of columns of M is linearly independent over Fp
iff it is independent over Q. There are infinitely many such primes, implying the following highly useful
result:
Theorem 5.4.3 (The finite field method). Let A ⊆ Rn be an integral hyperplane arrangement and q a power of a
large enough prime. Then χA (q) is the polynomial that counts points in the complement of Aq .
Example 5.4.4. Let G = ([n], E) be a simple graph and let AG be the corresponding graphic arrangement
in Rn . Note that AG reduces correctly over every finite field Fq (because graphic matroids are regular).
A point (x1 , . . . , xn ) ∈ Fnq can be regarded as the q-coloring of G that assigns color xi to vertex i. The
125
proper q-colorings are precisely the points of Fnq \ AG . The number of such colorings is pG (q) (the chromatic
polynomial of G evaluated at q). On the other hand, by Theorem 5.4.1, it is also the characteristic polynomial
χAG (q). Since pG (q) = χAG (q) for infinitely many q (namely, all integer prime powers), the polynomials
must be equal. In particular, by Zaslavsky’s theorems, the number of regions is pG (−1), and on the other
hand we know that regions of AG are in bijection with acyclic orientations of G (see Example 5.2.5), so we
now have a geometric proof of Theorem 4.5.2. ◀
Example 5.4.5. The Shi arrangement is the arrangement of n(n − 1) hyperplanes in Rn defined by
In other words, take the braid arrangement, clone it, and nudge each of the cloned hyperplanes a little bit in
the direction of the bigger coordinate. The Shi arrangement has rank n − 1 (every hyperplane in it contains
a line parallel to the all-ones vector), so we may project along that line to obtain the essentialization in Rn−1 .
Thus ess(Shi2 ) consists of two points on a line, while ess(Shi3 ) is shown below.
y =z+1
y=z
ess(Shi3 )
x=z+1
x=z
x=y x=y+1
(The number (n + 1)n−1 may look familiar; by Cayley’s formula, it is the number of spanning trees of
the complete graph Kn+1 . It also counts many other things of combinatorial interest, including parking
functions.)
The following proof is from [Sta07, §5.2]. By Theorem 5.4.3, it suffices to count the points in Fnq \ Shin for
a large enough prime q. Let x = (x1 , . . . , xn ) ∈ Fnq \ Shin . Draw a necklace with q beads labeled by the
elements 0, 1, . . . , q − 1 ∈ Fq , and for each k ∈ [n], put a big red k on the xk -th bead. For example, let n = 6
and q = 11. Then the necklace for x = (2, 5, 6, 10, 3, 7) is as follows:
126
0
10 1
4
9 2
1
8 5 3
6
7 4
3 2
6 5
The requirement that x avoids the hyperplanes xi = xj implies that the red numbers are all on different
beads. If we read the red numbers clockwise, starting at 1 and putting in a divider sign | for each bead
without a red number, we get
15 | 236 | | 4 |
which can be regarded as the ordered weak partition (or OWP)
that is, a (q − n)-tuple B1 , . . . , Bq−n , where the Bi are pairwise disjoint sets (possibly empty; that’s what the
“weak” means) whose union is [n], and 1 ∈ B1 . (We’ve omitted the divider corresponding to the bead just
counterclockwise of 1; stay tuned.)
Note that each block of Π(x) corresponds to a contiguous set of values among the coordinates of x. For
example, the block 236 occurs because the values 5,6,7 occur in coordinates x2 , x3 , x6 . In order to avoid the
hyperplanes xi = xj + 1 for i < j, each contiguous block of beads must have its red numbers in strictly
increasing order counterclockwise. (In particular the bead just counterclockwise of 1 must be unlabeled,
which is why we could omit that divider.)
To get a necklace from an OWP, write out each block in increasing order, with bars between successive
blocks.
Meanwhile, an OWP is given by a function f : [n] → [q − n], where f (i) is the index of the block containing i
(so f (1) = 1). There are (q − n)n−1 such things. Since there are q choices for the bead containing the red 1,
we obtain
Fnq \ Shin = q(q − n)n−1 = χShin (q).
This proves (5.9), and (5.10) follows from Zaslavsky’s theorems. ◀
We have seen that for a simple graph G = ([n], E), the chromatic polynomial pG (k) is precisely the char-
acteristic polynomial of the graphic arrangement AG . For some graphs, the chromatic polynomial factors
127
into linear terms over Z. For example, if G = Kn , then pG (k) = k(k − 1)(k − 2) · · · (k − n + 1), and if G is
a forest with n vertices and c components, then pG (k) = k c (k − 1)n−c . This property does not hold for all
graphs. For example, it is easy to work out that the chromatic polynomial of C4 (the cycle with four vertices
and four edges) is k 4 − 4k 3 + 6k 2 − 3k = k(k − 1)(k 2 − 3k + k), which does not factor further over Z. Is
there a structural condition on a graph or a central arrangement (or really, on a geometric lattice) that will
guarantee that its characteristic polynomial factors completely? It turns out that supersolvable geometric
lattices have this good property.
Definition 5.5.1. Let L be a ranked lattice. An element x ∈ L is a modular element if r(x) + r(y) =
r(x ∨ y) + r(x ∧ y) for every y ∈ L.
For example:
• By Theorem 1.5.6, a ranked lattice L is modular iff all elements are modular.
• The elements 0̂ and 1̂ are clearly modular in any lattice.
• If L is geometric, then every atom x is modular. Indeed, for y ∈ L, if y ≥ x, then y = x ∨ y and
x = x ∧ y, while if y ̸≥ x then y ∧ x = 0̂ and y ∨ x ⋗ y.
• The coatoms of a geometric lattice need not be modular. For example, let L = Πn , and recall that
Πn has rank function r(π) = n − |π|. Let x = 12|34, y = 13|24 ∈ Π4 . Then r(x) = r(y) = 2, but
r(x ∨ y) = r(1̂) = 3 and r(x ∧ y) = r(0̂) = 0. So x is not a modular element.
Proposition 5.5.2. The modular elements of Πn are exactly the partitions with at most one nonsingleton block.
X = {C ∈ σ : C ∩ B ̸= ∅}, Y = {C ∈ σ : C ∩ B = ∅}.
Then ( )
n o n o [
π ∧ σ = C ∩ B : C ∈ X ∪ {i} : i ̸∈ B , π∨σ = C ∪Y
C∈X
so
|π ∧ σ| + |π ∨ σ| = (|X| + n − |B|) + (1 + |Y |)
= (n − |B| + 1) + (|X| + |Y |) = |π| + |σ|,
For the converse, suppose B, C are nonsingleton blocks of π, with i, j ∈ B and k, ℓ ∈ C. Let σ be the
partition with exactly two nonsingleton blocks {i, k}, {j, ℓ}. Then r(σ) = 2 and r(π ∧ σ) = r(0̂) = 0, but
Modular elements are useful because they lead to factorizations of the characteristic polynomial of L.
Theorem 5.5.3. Let L be a geometric lattice of rank n, and let z ∈ L be a modular element. Then
X
χL (k) = χ[0̂,z] (k) µL (0̂, y)k n−r(z)−r(y) . (5.11)
y: y∧z=0̂
128
Here is a sketch of the proof; for the full details, see [Sta07, pp. 440–441]. We work in the dual Möbius alge-
bra A∗ (L) = A(L∗ ); that is, the vector space of C-linear combinations of elements of L, with multiplication
given by join (rather than meet as in §2.4). Thus the “algebraic” basis of A∗ (L) is
def X
{σy ≡ µ(y, x)x : y ∈ L}.
x: x≥y
for any z ∈ L. Second, for z, y, v ∈ L such that z is modular, v ≤ z, and y ∧ z = 0, one shows first
that z ∧ (v ∨ y) = v (by rank considerations) and then that rank(v ∨ y) = rank(v) + rank(y). Third, make
the substitutions v 7→ k rank z−rank v and y 7→ k n−rank y−rank z in the two sums on the RHS of (5.12). Since
vy = v ∨ y, the last observation implies that substituting x 7→ k n−rank x on the LHS preserves the product,
and the equation becomes (5.11).
P tell us anything new, because we already knew that k − 1 had to be a factor of χL (k),
This does not really
because χL (1) = x∈L µL (0̂, x) = 0. Also, the sum in the expression is not the characteristic polynomial of
a lattice.
On the other hand, if we have a modular coatom, then Theorem 5.5.3 is much more useful, since we can
identify an interesting linear factor and describe what is left after factoring it out.
Corollary 5.5.4. Let L be a geometric lattice, and let z ∈ L be a coatom that is a modular element. Then
If we are extremely lucky, then L will have a saturated chain of modular elements
0̂ = x0 ⋖ x1 ⋖ · · · ⋖ xn−1 ⋖ xn = 1̂.
In this case, we can apply Corollary 5.5.4 successively with z = xn−1 , z = xn−2 , . . . , z = x1 to split the
characteristic polynomial completely into linear factors:
where
129
Definition 5.5.5. A geometric lattice L is supersolvable if it has a modular chain, that is, a maximal chain
0̂ = x0 ⋖ x1 ⋖ · · · ⋖ xn = 1̂ such that every xi is a modular element. A central hyperplane arrangement A
is called supersolvable if L(A) is supersolvable.
Example 5.5.6. Every modular lattice is supersolvable, because every maximal chain is modular. In partic-
ular, the characteristic polynomial of every modular lattice splits into linear factors. ◀
Example 5.5.7. The partition lattice Πn (and therefore the associated hyperplane arrangement Brn ) is su-
persolvable by induction. Let z be the coatom with blocks [n − 1] and {n}, which is a modular element by
Proposition 5.5.2. There are n − 1 atoms a ̸≤ z, namely the partitions whose non-singleton block is {i, n} for
some i ∈ [n − 1], so we obtain
χΠn (k) = (k − n + 1)χΠn−1 (k)
and by induction
χΠn (k) = (k − 1)(k − 2) · · · (k − n + 1).
◀
Example 5.5.8. Let G = C4 (a cycle with four vertices and four edges), and let A = AG . Then L(A) is the
lattice of flats of the matroid U3 (4); i.e.,
L = {F ⊆ [4] : |F | =
̸ 3}
with r(F ) = min(|F |, 3). This lattice is not supersolvable, because no element at rank 2 is modular. For
example, let x = 12 and y = 34; then r(x) = r(y) = 2 but r(x ∨ y) = 3 and r(x ∧ y) = 0. (We have already
seen that the characteristic polynomial of L does not split.) ◀
Theorem 5.5.9. Let G = (V, E) be a simple graph. Then AG is supersolvable if and only if the vertices of G can be
ordered v1 , . . . , vn such that for every i > 1, the set
Ci := {vj : j ≤ i, vi vj ∈ E}
forms a clique in G.
Such an ordering is called a perfect elimination ordering. The proof of Theorem 5.5.9 is left as an exer-
cise (see Stanley, pp. 55–57). An equivalent condition is that G is a chordal graph: if C ⊆ G is a cycle of
length ≥ 4, then some pair of vertices that are not adjacent in C are in fact adjacent in G. This equivalence
is sometimes known as Dirac’s theorem. It is fairly easy to prove that supersolvable graphs are chordal, but
the converse is somewhat harder; see, e.g., [Wes96, pp. 224–226]. There are other graph-theoretic formu-
lations of this property; see, e.g., [Dir61]. See the recent paper [HS15] for much more about factoring the
characteristic polynomial of lattices in general.
If G satisfies the condition of Theorem 5.5.9, then we can see directly why its chromatic polynomial χ(G; k)
splits into linear factors. Consider what happens when we color the vertices in order. When we color vertex
vi , it has |Ci | neighbors that have already been colored, and they all have received different colors because
they form a clique. Therefore, there are k − |Ci | possible colors available for vi , and we see that
n
Y
χ(G; k) = (k − |Ci |).
i=1
One can also study complex hyperplane arrangements A ⊆ Cn . Since the hyperplanes of A have codimen-
sion 2 as real vector subspaces, the complement X = Cn \A is a connected topological space, but not simply
130
connected. Thus instead of counting regions, we should count holes, as expressed by the homology groups.
Brieskorn [Bri73] solved this problem completely:
Theorem 5.6.1 (Brieskorn [Bri73]). The homology groups Hi (X, Z) are free abelian, and the Poincáre polynomial
of X is the characteristic polynomial backwards:
n
X
rankZ Hi (X, Z)q i = (−q)n χL(A) (−1/q).
i=0
In a very famous paper, Orlik and Solomon [OS80] strengthened Brieskorn’s result by giving a presentation
of the cohomology ring H ∗ (X, Z) in terms of L(A), thereby proving that the cohomology is a combinatorial
invariant of A. (Brieskorn’s theorem says only that the additive structure of H ∗ (X, Z) is a combinatorial
invariant.) By the way, the homotopy type of X is not a combinatorial invariant; Rybnikov [Ryb11] con-
structed arrangements with isomorphic lattices of flats but different fundamental groups. There is much
more to say on this topic!
In another direction, one can study arrangements of subspaces of Rn or Cn that are not hyperplanes, i.e.,
have codimension greater than 1. This topic is much more difficult, in particular because one does not have
the nice combinatorial model of matroid theory in the background. A starting point is Björders survey
article [Bjö94].
Consider the two arrangements A1 , A2 ⊂ R2 shown in Figure 5.4. Their intersection posets are isomorphic,
so, by Zaslavsky’s theorems they have the same numbers of regions and bounded regions (this can of course
be checked directly). However, there is good reason not to consider the two arrangements isomorphic.
For example, both bounded regions in A1 are triangles, while A2 has a triangle and a trapezoid. Also,
the point H1 ∩ H2 ∩ H4 lies between the lines H3 and H5 in A1 , while it lies below both of them in A2 .
The intersection poset lacks the power to model geometric data like “between,” “below,” “triangle” and
“trapezoid.” Accordingly, we need to define a stronger combinatorial invariant.
H1 H1
H5 t s
H3 q r H3 q r
H4 H4
p p
s t
H5
A1 A2
H2 H2
First we fix notation. Let A = {H1 , . . . , Hn } be an essential hyperplane arrangement in Rd , with normal
vectors n1 , . . . , nn . For each i, let λi be an affine linear functional on Rn such that Hi = {x ∈ Rd : λi (x) = 0}.
(If ∩A = {⃗0} then we may define λi (x) = ni · x.)
131
The intersections of hyperplanes in A, together with its regions, decompose Rd as a polyhedral cell com-
plex: a disjoint union of polyhedra, each homeomorphic to Re for some e ≤ d (that’s what “cell” means),
such that the boundary of any cell is a union of other cells. We can encode each cell by recording whether
the linear functionals λ1 , . . . , λn are positive, negative or zero on it. Specifically, for k = (k1 , . . . , kn ) ∈
{+, −, 0}n , define a (possibly empty) subset of Rd by
λi (x) > 0 if ki = + ⇐⇒ i ∈ k+
F = F (k) = x ∈ Rd λi (x) < 0 if ki = − ⇐⇒ i ∈ k− .
λi (x) = 0 if ki = 0 ⇐⇒ i ∈ k0
This formula can be taken as the definition of k+ , k− , and k0 . A convenient shorthand (“digital notation”)
is to represent k by the list of digits i for which ki ̸= 0, placing a bar over the digits for which ki < 0. For
instance, k = 0 + −00 − +0 would be abbreviated 23̄6̄7; here k+ = {2, 7} and k− = {3, 6}.
If F ̸= ∅ then it is called a face of A, and k = k(F ) is the corresponding covector. The set of all faces is
denoted F (A). The poset Fˆ (A) = F (A) ∪ {0̂, 1̂}, ordered by containment of closures (F ≤ F ′ if F̄ ⊆ F̄ ′ ), is
a lattice, called the (big) face lattice1 of A. If A is central, then F (A) already has a unique minimal element
and we don’t add an extra one. For example, the big face lattice of Bool2 is shown in Figure 5.5.
2
1̂
Bool2 ∅ F (Bool2 )
2̄
Figure 5.5: The Boolean arrangement Bool2 and its big face lattice.
Combinatorially, the order relation in F (A) is given by k ≤ l if k+ ⊆ l+ and k− ⊆ l− . (This is very easy to
read off using digital notation.) The maximal covectors (or topes) are precisely those with no zeroes; they
correspond to the regions of A.
The big face lattice captures more of the geometry of A than the intersection poset; for instance, the two
arrangements A1 , A2 shown above have isomorphic intersection posets but non-isomorphic face lattices.
(This may be clear to you now; there are lots of possible explanations and we will see one soon.)
Example 5.7.1. An especially important example is the braid arrangement Brn (see Example 5.1.4), whose
faces have an explicit combinatorial description in terms of set compositions. If F is a face, then F lies
either below, above, or on each hyperplane Hij — i.e., either xi < xj , xi = xj , or xi > xj holds on F —
and this data describes F exactly. In fact, we can record F by a set composition of [n], i.e., an ordered list A
of nonempty sets A1 | . . . |Ak whose disjoint union is [n]. (We write A |= [n] for short.) For example, the set
composition
569 | 3 | 14 | 28 | 7
1 That is, the big lattice of faces, not the lattice of big faces.
132
x=z
x<z
1|3|2
x>z
3|1|2 13|2 1|23 1|2|3 1|2|3 1|3|2 2|1|3 3|1|2 2|3|1 3|2|1
123 x<y
3|12 12|3 x=y 1|23 12|3 13|2 2|13 3|12 23|1
x>y
y>z
2|3|1
y<z
y=z
Figure 5.6: Br3 and its big face lattice (the lattice of set compositions).
Note that the number of blocks of A (in this case, 5) equals dim FA , since that is the number of free coor-
dinates. In fact, FA is linearly equivalent to a maximal region of Br5 , say the principal region, under the
linear transformation R5 → R9 given by (a, b, c, d, e) 7→ (c, d, b, c, a, a, e, d, a); in particular it is a simplicial
polyhedron. In the extreme case that dim A = n, the set composition has only singleton parts, hence is
equivalent to a permutation (this confirms what we already know, that Brn has n! regions).
The correspondence between faces of Brn and set compositions A |= [n] is a bijection. In fact, the big face
lattice of Brn is isomorphic to the lattice of set compositions ordered by refinement; see Figure 5.6. ◀
More generally, consider a system of linear equalities and inequalities of the form xi = xj and xi < xj . If
such a system is consistent, it gives rise to a nonempty polyhedron that is a convex union of faces of Br9 .
Such a system can be described by a preposet, which is a relation < on [n] that is reflexive, transitive, but
not necessarily antisymmetric (compare Defn. 1.1.1). In other words, x ≤ y and y ≤ x does not imply x = y.
This relation has a Hasse diagram, just like a poset, except that multiple elements of the ground set can be
put in the same “box” (whenever there is a failure of antisymmetry). For example, the system
3 15 268
4 9
133
and this gives rise to a 6-dimensional convex polyhedron P consisting of faces of Br9 . (Each box in the
Hasse diagram represents a coordinate that can vary (locally) freely, which is why the dimension is 6.)
The maximal faces in P correspond to the linear extensions of the preposet, expressed as set compositions:
15|3|4|9|268|7, 4|9|268|3|15|7, etc. For more on the “cone/preposet dictionary”, see [PRW08].
Oriented matroids are a vast topic; these notes just scratch the surface. The canonical resource is the
book [BLVS+ 99]; an excellent free source is Reiner’s lecture notes [Rei] and another good brief reference
is [RGZ97].
Consider the linear forms λi that were used in representing each face by a covector. Recall that specifying λi
is equivalent to specifying a normal vector ni to the hyperplane Hi (with λi (x) = ni · x). As we know, the
vectors ni represent a matroid whose lattice of flats is precisely L(A). Scaling ni (equivalently, λi ) by a
nonzero constant c ∈ R has no effect on the matroid represented by the ni ’s, but what does it do to the
covectors? If c > 0, then nothing happens, but if c < 0, then we have to switch + and − signs in the ith
position of every covector. So, in order to figure out the covectors, we need not just the normal vectors ni ,
but an orientation for each one — hence the term “oriented matroid”. Equivalently, for each hyperplane Hi ,
we are designating one of the two corresponding halfspaces (i.e., connected components of Rd \ Hi ) as
positive and the other as negative.
See Figure 5.7 for examples. (The normal vectors all have positive z-coordinate, so “above” means “above.”)
For instance, the trapezoidal bounded region in A2 has covector ++++− because it lies above hyperplanes
H1 , H2 , H3 , H4 but below H5 . Its top side has covector + + + + 0, its bottom + + 0 + −, etc.
H1 H1
−++++ +++++ +−+++
H5 s
t
−++++ +++++ +−+++ −+++− ++++− +−++−
H3 q r
H3 q r
++−++ ++−+−
−+−++ +−−++ −+−+− +−−+−
H4 H4
−+−−+
p +−−−+
p
s −−−++
t
H5 −+−−− −−−−− +−−−−
A1 A2
H2 H2
Proposition 5.8.1. Suppose that no two hyperplanes in A are parallel. Then the maximal covectors whose negatives
are also covectors are precisely those that correspond to relatively-unbounded faces. In particular, A is central if and
134
only if every negative of a covector is a covector.
Suppose R is an unbounded region, with k the corresponding covector. Fix a point x ∈ R and choose a
direction v in which R is unbounded. By perturbing v slightly, we can assume that v is not orthogonal
to any normal vector ni for which ki = 0. (This perturbation step is where we use the assumption that
no two hyperplanes are parallel.) In other words, if we walk in the direction of v then the values of λi
increase without bound, decrease without bound, or remain zero according as i belongs to k+ , k− , or k0 .
But then if we walk in the direction of −ni , then “increase” and “decrease” are reversed. Therefore, walking
sufficiently far in that direction arrives in an (unbounded) region with covector −k.
Conversely, suppose that k and −k are covectors of regions R and S. Pick points x ∈ R and y ∈ S and
consider the line ℓ joining x and y. The functionals λi are identically zero on ℓ for i ∈ k0 = (−k)0 , but
otherwise increase or decrease (necessarily without bound). Therefore the ray pointing from x away from
y (resp., from y away from x) is contained in R (resp., S). It follows that both R and S are unbounded.
The second assertion now follows from Corollary 5.3.10. WHICH IS FALSE
(It would be nice to modify the statement to handle the case that A has parallel hyperplanes. Here the
conclusion fails, since for example in A1 or A2 above, every ray in the region with covector + − − − + is
horizontal, hence orthogonal to the normals to H3 , H4 .H5 , so the functionals λ3 , λ4 are constant and positive
— hence do not become negative upon walking in the other direction; the “opposite” unbounded region
has covector −+−−+. It is still true that any pair of opposite covectors correspond to opposite unbounded
regions, but I think this condition holds only for unbounded regions that contain more than one direction’s
worth of rays.)
Just like circuits, bases, etc., of a matroid, oriented matroid covectors can be axiomatized purely combina-
torially. First some preliminaries. For k, l ∈ {+, 0, −}n , define the composition k ◦ l by
(
ki if ki ̸= 0,
(k ◦ l)i =
li if ki = 0.
The axioms are as follows [RGZ97, §7.2.1]: a collection K ⊆ {+, −, 0}n is a covector system if for all
k, l ∈ K :
(K1) ⃗0 = (0, 0, . . . , 0) ∈ K ;
(K2) −k ∈ K ;
(K3) k◦l∈K ;
(K4) If i ∈ S(k, l) then there exists m ∈ K with (a) mi = 0 and (b) mj = (k ◦ l)j for j ∈ [n] \ S(k, l).
Note that (K1) and (K2) are really properties of central hyperplane arrangements. However, any non-central
arrangement A can be turned into a central one by coning (see Definition 5.1.8), and if K (A) is the set of
covectors of A then
K (cA) = {(k, +) : k ∈ K (A)} ∪ {(−k, −) : k ∈ K (A)} ∪ {⃗0}
and by the way, K (A) = {k : (k, +) ∈ K (cA)}.
135
5.8.2 Oriented matroid circuits
The cones over the arrangements A1 and A2 (not including the new hyperplane introduced in coning) are
central, essential arrangements in R3 , whose matroids of normals can be represented respectively by the
matrices
Evidently the matroids represented by X1 and X2 are isomorphic, with circuit system {124, 345, 1235}.
However, they are not isomorphic as oriented matroids. The minimal linear dependencies realizing the
circuits in each case are
An oriented circuit keeps track not just of minimal linear dependencies, but of how to orient the vectors
in the circuit so that all the signs are positive. Thus 124̄ is a oriented circuit in both cases. However, in the
first case 34̄5 is a circuit, while in the second it is 34̄5̄. Note that if c is a circuit then so is −c, where, e.g.,
−124̄ = 1̄2̄4. In summary, the oriented circuit systems for CA1 and CA2 are respectively
Oriented circuits are minimal obstructions to covector-ness. For example, 124̄ is a circuit of A1 because the
linear functionals defining its hyperplanes satisfy λ1 + λ2 − 2λ4 = 0. But if a covector of A1 contains 124̄,
then any point in the corresponding face of A would have λ1 , λ2 , −λ4 all positive, which is impossible.
(OC1) ⃗0 ̸∈ C⃗.
(OC2) −c ∈ C⃗.
(OC3) Either c+ ̸⊆ c′+ or c− ̸⊆ c′− .
(OC4) If ci = + and c′i = −, then there exists d ∈ C⃗ such that (a) di = 0 and (b) for all j ̸= i, d+ ⊆ c+ ∪ c′+
and d− ⊆ c− ∪ c′− .
Again, the idea is to record not just the linearly dependent subsets of a set {λi , . . . , λn } of linear forms, but
also the sign patterns of the corresponding linear dependences (“syzygies”). The first two are elementary:
(OC1) says that the empty set is linearly independent and (OC2) says that multiplying any syzygy by −1
gives a syzygy. Condition (OC3) must hold if we want circuits to record signed syzygies with minimal
support, as for circuits in an unoriented matroid,
136
(OC4) is the oriented version of circuit exchange. Suppose that we have two syzygies
n
X n
X
γj λ j = γj′ λj = 0
j=1 j=1
with γi > 0 and γi′ < 0 for some i. Multiplying by positive scalars if necessary (hence not changing the sign
patterns), we may assume that γi = −γi′ . Then adding the two syzygies gives
n
X
δj λj = 0,
j=1
where δj = γj + γj′ . In particular, δi = 0, and δj is positive (resp., negative) if and only if at least one of γj , γj′
is positive (resp., negative).
Remark 5.8.3. If C⃗ is an oriented circuit system, then C = {c+ ∪ c− : c ∈ C⃗} is a circuit system for an
ordinary matroid with ground set [n]. (I.e., just erase all the bars.) This is called the underlying matroid of
the oriented matroid with circuit system C⃗.
As in the unoriented setting, the circuits of an oriented matroid represent minimal obstructions to being a
covector. That is, every real hyperplane arrangement A gives rise to an oriented circuit system C⃗ such that
if k is a covector of A and c is a circuit, then it is not the case that k+ ⊇ c+ and k− ⊇ c− .
More generally, one can construct an oriented matroid from any real pseudosphere arrangement, or collection
of homotopy (d − 1)-spheres embedded in Rn such that the intersection of the closures of the spheres in any
subcollection is either connected or empty — i.e., a thing like this:
Again this arrangement gives rise to a cellular decomposition of Rn , and each cell corresponds to a covector
which describes whether the cell is inside, outside, or on each pseudocircle.
In fact, the Topological Representation Theorem of Folkman and Lawrence (1978) says that every combi-
natorial oriented matroid can be represented by such a pseudosphere arrangement. However, there exist
oriented matroids that cannot be represented as hyperplane arrangements. For example, recall the con-
struction of the non-Pappus matroid (Example 3.5.7). If we bend the line xyz a little so that it meets x and y
but not z (and no other points), the result is a pseudoline arrangement whose oriented matroid M cannot
be represented by means of a line arrangement.
137
5.8.3 Oriented matroids from graphs
Recall (§3.3) that every graph G = (V, E) gives rise to a graphic matroid M (G) with ground set E. Corre-
spondingly, every directed graph G⃗ gives rise to an oriented matroid, whose circuit system C⃗ is the family of
oriented cycles. This is best shown by an example.
2
Oriented circuits
For example, 135̄ is a circuit because the clockwise orientation of the northwest triangle in G includes edges
1 and 3 forward, and edge 5 backward. In fact, this circuit system is identical to the circuit system C1 seen
previously. More generally, for every oriented graph G, ⃗ the signed set system C⃗ formed in this way satisfies
the axioms of Definition 5.8.2. To understand axiom (4) of that definition, suppose e is an edge that occurs
forward in c and backward in c′ . Then c − e and c′ − e are paths between the two endpoints of e, with
opposite starting and ending points, so when concatenated, they form an closed walk in G, ⃗ which must
contain an oriented cycle.
Reversing the orientation of edge e corresponds to interchanging e and ē in the circuit system; this is called
a reorientation. For example, reversing edge 5 produces the previously seen oriented circuit system C2 .
An oriented matroid is called acyclic if every circuit has at least one barred and at least one unbarred
⃗ having no directed cycles (i.e., being an acyclic orientation of its underlying
element; this is equivalent to G
graph G). In fact, for any ordinary unoriented matroid M , one can define an orientation of M as an oriented
matroid whose underlying matroid is M ; the number of acyclic orientations is TM (2, 0) [Rei, §3.1.6, p.29],
just as for graphs.
The covectors of the circuit system for a directed graph are in fact the faces of the (essentialization of) the
graphic arrangement associated to G,⃗ in which the orientation of each edge determines the orientation of
⃗ then the hyperplane xi = xj is assigned the normal
the corresponding normal vector — if ı⃗ȷ is an edge in G
vector ei − ej . The maximal covectors are precisely the regions of the graphic arrangement.
5.9 Exercises
Problem 5.1. Let m > n, and let A be the arrangement of m affine hyperplanes in general position in
Rn . Here “general position” means that every k of the hyperplanes intersect in an affine linear space of
dimension n − k; if k > n then the intersection is empty. (Compare Example 5.3.9, where the hyperplanes
are linear.) Calculate χA (k), r(A), and b(A).
Solution: The arrangement A is certainly essential. Its intersection poset is a Boolean lattice Boolm that has
been truncated above rank n, i.e.,
L(A) = {S ⊆ [m] : |S| ≤ n},
138
with Möbius function µ(0̂, S) = (−1)|S| . Its characteristic polynomial is therefore
n
X X X m n−s
χA (k) = µ(0̂, x) k dim x = (−1)|S| k n−|S| = (−1)s k .
s=0
s
x∈L(A) S⊆[m]: |S|≤n
To prove the last equality, consider the set S of all subsets of [m] of size ≤ n. Within this set, there is a
bijection
ϕ : {S ∈ S : m ∈ S} → {S ∈ S : m ̸∈ S, |S| < n}
given by S 7→ S \ {m}. Since each pair {S, ϕ(S)} has cardinalities that differ by 1, their net contribution to
the sum
n
n−s m
X X
(−1) = (−1)n−|S|
s=0
s
S∈S
is zero. Therefore
m−1
X X X
(−1)n−|S| = (−1)n−|S| = (−1)n−|S| =
n
S∈S S∈S : m̸∈S, |S|=n S∈([m−1]
n )
m2 + m + 2 m2 − 3m + 2
m m m m m m
r(A) = + + = , b(A) = − + = ,
0 1 2 2 0 1 2 2
confirming Example 5.2.3.
Problem 5.2. (Stanley, HA, 2.5) Let G be a graph on n vertices, let AG be its graphic arrangement in Rn ,
and let BG = Booln ∪ AG . (That is, B consists of the coordinate hyperplanes xi = 0 in Rn together with the
hyperplanes xi = xj for all edges ij of G.) Calculate χBG (q) in terms of χAG (q).
Solution: For any prime power q, the points of Fnq \ (BG ⊗ Fq ) can be interpreted as proper colorings of G
with colors chosen from Fq \ {0}. Since the characteristic polynomial of AG is the chromatic polynomial
of G, we conclude that
χBG (q) = χAG (q − 1).
Problem 5.3. (Stanley, EC2, 3.115) Determine the characteristic polynomial and the number of regions of
the type B braid arrangement and the type D braid arrangement Bn , Dn ⊂ Rn , which are defined by
Solution: Again, we count the number of points x = (x1 , . . . , xn ) of the complements of Bn and Dn in Fnq ,
where q is large enough so that the arrangements reduce correctly. For Bn , there are
139
• q − 1 choices for x1 — any element of Fq \ {0};
• q − 3 choices for x2 — any element of Fq \ {0, x1 , −x1 };
• ...
• q − 2n + 1 choices for xn .
Therefore,
n
Y
χBn (q) = (q − 2k + 1).
k=1
Meanwhile, aQ point x ∈ Fnq \ Dn either has no zero coordinates (when it is in the complement of Bn , so
n
that there are k=1 (q − 2k + 1) possibilities) or exactly one zero coordinate. In the second possibility, there
are n possibilities for the zero coordinate and, by a similar argument to the first case, there are (q − 1)(q −
3) · · · (q − 2n + 3) choices for the remaining n − 1 coordinates. Therefore,
n
Y n−1
Y
χDn (q) = (q − 2k + 1) + n (q − 2k + 1)
k=1 k=1
n−1
Y
= (q − n + 1) (q − 2k + 1)
k=1
= (q − n + 1)(q − 1)(q − 3) · · · (q − 2n + 3).
Problem 5.4 (Stanley [Sta07], Exercise 5.9(a)). Find the characteristic polynomial and number of regions of
the arrangement An ⊆ Rn with hyperplanes xi = 0, xi = xj , and xi = 2xj , for all 1 ≤ i ̸= j ≤ n.
has n coordinates that correspond to n non-adjacent beads. So we first count the number of ways to choose
such a set of beads including the bead 1. Let •◦ denote including a bead but excluding the next bead, and
let ◦ denote excluding a bead. Once we have chosen to include bead 1 and not bead 2, we need to fill out
the rest of the necklace with n − 1 copies of •◦ and therefore (q −1) − 2 − 2(n − 1) = q − 2n − 1 copies of
◦. The number of ways we can do this is evidently n−1+q−2n−1 n−1 = q−n−2
n−1 . Any bead is as likely to be
used as any other, so the probability that bead 1 is used is n(q − 1); therefore, the number of choices for the
coordinates of x is
q−1 q−n−2 q−1 (q − n − 2)! (q − 1)(q − n − 2)!
= · =
n n−1 n (n − 1)!(q − 2n − 1)! n!(q − 2n − 1)!
and once we know the set of coordinates, there are n! ways to assign them to x1 , . . . , xn , giving
(q − 1)(q − n − 2)!
χA (q) = = (q − 1)(q − n − 2)(q − n − 3) · · · (q − 2n)
(q − 2n − 1)!
140
and therefore
2(2n + 1)!
r(A) = |χA (−1)| = .
(n + 2)!
Stanley says that proving this combinatorially (i.e., by finding a bijection between the regions of A and
some set whose cardinality is “obviously” 2(2n+1)!
(n+2)! ) is open. Interestingly, the same argument works upon
replacing the hyperplanes xi = 2xj with xi = pxj for any prime p (provided Artin’s conjecture holds!). It
does not work upon replacing p with a composite number, although the result may still be true.
Problem 5.5. Recall that each permutation w = (w1 , . . . , wn ) ∈ Sn corresponds to a region of the braid
arrangement Brn , namely the open cone Cw = {(x1 , . . . , xn ) ∈ Rn : xw1 < xw2 < · · · < xwn }. Denote its
closure by Cw . For any set W ⊆ Sn , consider the closed fan
[
F (W ) = Cw = {(x1 , . . . , xn ) ∈ Rn : xw1 ≤ · · · ≤ xwn for some w ∈ W }.
w∈W
Prove that F (W ) is a convex set if and only if W is the set of linear extensions of some poset P on [n]. (A
linear extension of P is a total ordering ≺ consistent with the ordering of P , i.e., if x <P y then x ≺ y.)
Solution: To be written
Problem 5.6. The runners in a sprint are seeded 1, . . . , n (stronger runners are assigned higher numbers).
To even the playing field, the rules specify that you earn one point for each higher-ranked opponent you
beat, and one point for each lower-ranked opponent you beat by at least one second. (If a higher-ranked
runner beats a lower-ranked runner by less than 1 second, no one gets the point for that matchup.) Let si
be the number of points scored by the ith player and let s = (s1 , . . . , sn ) be the score vector.
(a) Show that the possible score vectors are in bijection with the regions of the Shi arrangement.
(b) Work out all possible score vectors in the cases of 2 and 3 players. Conjecture a necessary and sufficient
condition for (s1 , . . . , sn ) to be a possible score vector for n players. Prove it if you can.
Solution: To be written
Solution: To be written
141
Chapter 6
Simplicial Complexes
The canonical references for this material are [Sta96], [BH93, Ch. 5]. See also [MS05] (for the combinatorics
and algebra) and [Hat02] (for the topology).
Definition 6.1.1. Let V be a finite set of vertices. An (abstract) simplicial complex ∆ on V is a nonempty
family of subsets of V with the property that if σ ∈ ∆ and τ ⊆ σ, then τ ∈ ∆. Equivalently, ∆ is an order
ideal in the Boolean lattice 2V . The elements of ∆ are called its faces or simplices. A face that is maximal
with respect to inclusion is called a facet.
The dimension of a face σ is dim σ = |σ| − 1. A face of dimension k is a k-face or k-simplex. The dimension
of a non-void simplicial complex ∆ is dim ∆ = max{dim σ : σ ∈ ∆}. (Sometimes we write ∆d−1 to indicate
that dim ∆ = d − 1; this is a common convention since then d is the maximum number of vertices in a face.)
A complex is pure if all its facets have the same dimension.
The simplest simplicial complexes are the void complex ∆ = ∅ (which is often excluded from consideration)
and the irrelevant complex ∆ = {∅}. In some contexts, there is the additional requirement that every
singleton subset of V is a face (since if v ∈ V and {v} ̸∈ ∆, then v ̸∈ σ for all σ ∈ ∆, so you might as well
replace V with V \ {v}). A simplicial complex with a single facet is also called a simplex.
The set of facets of a complex is the unique minimal set of generators for it.
Simplicial complexes are combinatorial models for compact topological spaces. The vertices V = [n] can
be regarded as the points e1 , . . . , en ∈ Rn , and a simplex σ = {v1 , . . . , vr } is then the convex hull of the
corresponding points:
142
For example, faces of sizes 1, 2, and 3 correspond respectively to vertices, line segments, and triangles. (This
explains why dim σ = |σ| − 1.) Taking {ei } to be the standard basis of Rn gives the standard geometric
realization |∆| of ∆: [
|∆| = conv{ei : i ∈ σ}.
σ∈∆
It is usually possible to realize ∆ geometrically in a space of much smaller dimension. For example, every
graph can be realized in R3 , and planar graphs can be realized in R2 . It is common to draw geometric
pictures of simplicial complexes, just as we draw pictures of graphs. We sometimes use the notation |∆| to
denote any old geometric realization (i.e., any topological space homeomorphic to the standard geometric
realization). Typically, it is easiest to ignore the distinction between ∆ and |∆|; if we want to be specific we
will use terminology like “geometric realization of ∆” or “face poset of ∆”. A triangulation of a topological
space X is a simplicial complex whose geometric realization is homeomorphic to X.
Figure 6.1 shows geometric realizations of the simplicial complexes ∆1 = ⟨124, 23, 24, 34⟩ and ∆2 = ⟨12, 14, 23, 24, 34⟩.
2 2
1 3 1 3
4 4
∆1 ∆2
The filled-in triangle indicates that 124 is a face of ∆1 , but not of ∆2 . Note that ∆2 is the subcomplex of ∆1
consisting of all faces of dimensions ≤ 1 — that is, it is the 1-skeleton of ∆1 .
The link can be thought of as “what you see if you stand in σ in look outward”; for example, if ∆ is a
triangulation of a (d − 1)-dimensional manifold, then the link of every vertex is a (d − 2)-sphere, and
more generally the link of every k-dimensional face is a (d − k − 2)-sphere.
4. The join of two complexes ∆, ∆′ on disjoint vertex sets is
∆ ∗ ∆′ = {σ ∪ σ ′ : σ ∈ ∆, σ ′ ∈ ∆′ }.
Combinatorially, join behaves like a product; for example, it is multiplicative on f -vectors. On the
other hand, it is not a product in the topological sense: [∆ ∗ ∆′ ] is not homeomorphic to [∆] × [∆′ ].
143
And here is the basic numerical invariant of a simplicial complex.
Definition 6.1.2. Let ∆d−1 be a simplicial complex. The f -vector of ∆ is (f−1 , f0 , f1 , . . . , fd−1 ), where fi =
fi (∆) is the number of faces of dimension i. The term f−1 is often omitted, because f−1 = 1 unless ∆ is the
void complex. The f -polynomial is the generating function for the nonnegative f -numbers (essentially the
rank-generating function of ∆ as a poset):
Example 6.1.3. Let P be a finite poset and let ∆(P ) be the set of chains in P . Every subset of a chain is a
chain, so ∆(P ) is a simplicial complex, called the order complex of P . The minimal nonfaces of ∆(P ) are
precisely the pairs of incomparable elements of P ; in particular every minimal nonface has size two, which
is to say that ∆(P ) is a flag complex. Note that ∆(P ) is pure if and only if P is ranked.
If P itself is the set of faces of a simplicial complex ∆, then ∆(P (∆)) is the barycentric subdivision of that
complex. Combinatorially, the vertices of Sd(∆) correspond to the faces of ∆; a collection of vertices of
Sd(∆) forms a face if the corresponding faces of ∆ are a chain in its face poset. Topologically, Sd(∆) can
be constructed by drawing a vertex in the middle of each face of ∆ and connecting them — this is best
illustrated by a picture.
∆ Sd(∆)
Each vertex (black, red, blue) of Sd(∆) corresponds to a (vertex, edge, triangle) face of ∆. Note that barycen-
tric subdivision does not change the topological space itself, only the triangulation of it. ◀
Simplicial complexes are models of topological spaces, and combinatorialists use tools from algebraic topol-
ogy to study them, in particular the machinery of simplicial homology. Here we give a “user’s guide” to the
subject that assumes as little topology background as possible. Readers familiar with the subject will know
that I am leaving many things out. For a full theoretical treatment, I recommend Chapter 2 of Hatcher
[Hat02].
Let ∆ be a simplicial complex on vertex set [n]. The kth simplicial chain group of ∆ over a field1 , say R,
is the vector space Ck (∆) of formal linear combinations of k-simplices in ∆. Thus dim Ck (∆) = fk (∆). The
1 More generally, this could be any commutative ring, but let’s keep things simple for the moment.
144
elements of Ck (∆) are called k-chains. The (simplicial) boundary map ∂k : Ck (∆) → Ck−1 (∆) is defined
as follows: if σ = {v0 , . . . , vk } is a k-face, with 1 ≤ v0 < · · · < vk ≤ n, then
k
X
∂k [σ] = (−1)i [v0 , . . . , vbi , . . . , vk ] ∈ Ck−1 (∆)
i=0
where the hat denotes removal. The map is then extended linearly to all of Ck (∆).
The entire collection of data {Ck (∆, ∂k )} is called the simplicial chain complex of ∆. For example, if
∆ = ⟨123, 14, 24⟩, then the simplicial chain complex is as follows:
∂ ∂ ∂
C2 = R1 −−−−−2−−−→ C1 = R5 −−−−−−−−−−−−−
1
−−−−−−−−−−→ C0 = R4 −−−−−−−−0−−−−−−→ C−1 = R
123 12 13 14 23 24 1 2 3 4
12 1 1 1 1 1 0 0 ∅ [1 1 1 1]
13 −1 2 −1 0 0 1 1
14 0 3 0 −1 0 −1 0
−1 −1
23 1 4 0 0 0
24 0
The fundamental fact about boundary maps is that ∂k ◦∂k+1 for all k, a fact that is frequently written without
subscripts:
∂ 2 = 0.
(This can be checked directly from the definition of ∂, and is a calculation that everyone should do for
themselves once.) This is precisely what the term “chain complex” means in algebra.
An equivalent condition is that ker ∂k ⊇ im ∂k+1 for all k. In particular, we can define the reduced simplicial
homology groups2
H̃k (∆) = ker ∂k / im ∂k+1 .
The H̃k (∆) are just R-vector spaces, so they can be described up to isomorphism by their dimensions3 ,
which are called the Betti numbers βk (∆). They can be calculated using the rank-nullity formula: in general
βk (∆) = dim H̃k (∆) = dim ker ∂k − dim im ∂k+1 = fk − rank ∂k − rank ∂k+1 .
These numbers turn out to carry topological information about the space |∆|. In fact, they depend only
on the homotopy type of the space |∆|. This is a fundamental theorem in topology whose proof is far too
2 The unreduced homology groups H (∆) are defined by deleting C
k −1 (∆) from the simplicial chain complex. This results in an ex-
tra summand of R in H0 (∆) and has no effect elsewhere. Broadly speaking, reduced homology arises more naturally in combinatorics
and unreduced homology is more natural in topology, but the information is equivalent.
3 This would not be true if we replaced R with a ring that was not a field. Actually, the most information is available over Z. In that
case βk (∆) can still be obtained as the rank of the free part if H̃k (∆), but there also may be a torsion part.
145
elaborate to give here,4 but provides a crucial tool for studying simplicial complexes: we can now ask how
the topology of ∆ affects its combinatorics. To begin with, the groups H̃k (∆) do not depend on the choice
of labeling of vertices and are invariant under retriangulation.
A complex all of whose homology groups vanish is called acyclic. For example, if |∆| is contractible then
∆ is acyclic over every ring. If ∆ ∼
= Sd (i.e., |∆| is a d-dimensional sphere), then
(
R if k = d,
H̃k (∆) ∼= (6.1)
0 if k < d.
The second equality here is called the Euler-Poincaré theorem; despite the fancy name, it is easy to prove
using little more than the rank-nullity theorem of linear algebra (Problem 6.10). The Euler characteristic is
the single most important numerical invariant of ∆. Many combinatorial invariants can be computed by
identifying them as the Euler characteristic of a simplicial complex whose topology is known, often one
that is acyclic (χ̃ = 0), a sphere of dimension d (χ̃ = (−1)d ), or a wedge of spheres.
Observe that
X X
χ̃(∆) = (−1)dim σ + (−1)dim σ
σ∈∆: e̸∈σ σ∈∆: e∈σ
X X
dim σ
= (−1) + (−1)1+dim τ
σ∈del∆ (e) τ ∈link∆ (e)
Definition 6.3.1. Let ∆ be a simplicial complex on vertex set [n]. Its Stanley-Reisner ideal in R is
invariants but are well-nigh impossible to work with directly; one then shows that repeatedly barycentrically subdividing a space
allows us to approximate singular homology by simplicial homology sufficiently accurately — but on the other hand subdivision also
preserves simplicial homology, so we can have the best of both worlds. See [Hat02, §2.1] for the full story. Or take my MATH 821
class!
146
Example 6.3.2. Let ∆1 and ∆2 be the complexes of Figure 6.1. Abbreviating w, x, y, z = x1 , x2 , x3 , x4 , the
Stanley-Reisner ideal of ∆1 is
I∆1 = ⟨wxyz, wxy, wyz, xyz, wy⟩ = ⟨xyz, wy⟩.
Note that the minimal generators of I∆ are the minimal nonfaces of ∆. Similarly,
I∆2 = ⟨wxz, xyz, wy⟩.
If ∆ is the simplex on [n] then it has no nonfaces, so I∆ is the zero ideal and k[∆] = k[x1 , . . . , xn ]. In general,
the more faces ∆ has, the bigger its Stanley-Reisner ring is. ◀
Since ∆ is a simplicial complex, the monomials in I∆ are exactly those whose support is not a face of ∆.
Therefore, the monomials supported on a face of ∆ are a natural vector space basis for the graded ring
k[∆]. Its Hilbert series can be calculated by counting these monomials:
def X i X X
Hilb(k[∆d−1 ], q) ≡ q dimk (k[∆])i = q deg(µ)
i≥0 σ∈∆ monomials µ:
supp µ=σ
X q |σ|
=
1−q
σ∈∆
d
X d
X
d i fi−1 q i (1 − q)d−i hi q i
X q i=0 i=0
= fi−1 = =
i=0
1−q (1 − q)d (1 − q)d
The numerator of this rational expression is a polynomial in q, called the h-polynomial of ∆ and written
h∆ (q), and its list of coefficients (h0 , h1 , . . . , HD ) is called the h-vector of ∆. Clearing denominators and
applying the binomial theorem yields a formula for the h-numbers in terms of the f -numbers:
d d d d−i
d−i
X X X X
hi q i = fi−1 q i (1 − q)d−i = fi−1 q i (−1)j q j
i=0 i=0 i=0 j=0
j
d Xd−i
d−i
X
= (−1)j q i+j fi−1
i=0 j=0
j
and now extracting the q k coefficient (i.e., the summand in the second sum with j = k − i) yields
k
d−i
X
hk (∆) = (−1)k−i fi−1 (∆). (6.4)
i=0
k − i
where dim ∆ = d − 1. (Note that the upper limit of summation might as well be k instead of d, since the
binomial coefficient in the summand vanishes for i > k.) These equations can be solved to give the f ’s in
terms of the h’s.
i
d−k
X
fi−1 (∆) = hk (∆). (6.5)
i−k
k=0
So the f -vector and h-vector contain the same information about ∆. On the level of generating functions,
the conversions look like this [BH93, p. 213]:
X X
hi q i = fi−1 q i (1 − q)d−i , (6.6)
i i
X X
i
fi q = hi q i−1 (1 + q)d−i . (6.7)
i i
147
The equalities (6.4) and (6.5) can be obtained by applying the binomial theorem to the right-hand sides
of (6.6) and (6.7) and equating coefficients. Note that it is most convenient simply to sum over all i ∈ Z.
Let’s go back to the formula for the Hilbert series in terms of the h-vector, namely
d
X
hi q i
h∆ (q)
Hilb(k[∆], q) = = i=0 .
(1 − q)d (1 − q)d
Note that 1/(1 − q)d is just the Hilbert series of the polynomial ring k[x1 , . . . , xn ]. More generally, if R is any
graded ring and x is an indeterminate of degree 1, then
Hilb(R, q)
Hilb(R[x], q) = .
1−q
These observations suggest that we should be able to regard the Stanley-Reisner ring k[∆] as a polynomial
ring in d variables over a base ring S whose Hilbert series is the polynomial h∆ (q). In particular S would
have to be a finite-dimensional vector space and hi the dimension of its ith graded piece. Also, we should
be able to recover S by quotienting out by d linear forms, each of which would remove a factor of 1/(1 − q)
from the Hilbert series. A ring for which this all works is called a Cohen-Macaulay (CM) ring, and a
Cohen-Maculay simplicial complex is one whose Stanley-Reisner ring is CM.
Example 6.3.3. The bowtie complex is the pure 2-dimensional complex ∆ = ⟨123, 145⟩ shown below, with
f -vector (1, 5, 6, 2). Therefore, by (6.6), the h-polynomial is
X
hi q i = 1q 0 (1 − q)3 + 5q 1 (1 − q)2 + 6q 2 (1 − q)1 + 2q 3 (1 − q)0 = 1 + 2q − q 2
i
2 4
1
3 5
This complex cannot possibly be CM, because if k[∆] were the ring of polynomials in two indeterminates
over a subring S, then the Hilbert series of S would have to be 1 + 2q − q 2 — in particular, the degree-2
graded piece of S would be a vector space of dimension −1, which is absurd. ◀
In particular, the h-numbers of a Cohen-Macaulay simplicial complex are all nonnegative, suggesting that
they should count something (that is, something more combinatorial then the dimensions of the subring
over which k[∆] is a polynomial ring).
148
6.4 Shellable and Cohen-Macaulay simplicial complexes
Here is an important special class of complexes where the h-numbers have a direct combinatorial interpre-
tation.
Definition 6.4.1. A pure simplicial complex ∆d−1 is shellable if its facets can be ordered F1 , . . . , Fn such
that any of the following conditions are satisfied:
1. For every i ∈ [n], the set Ψi = ⟨Fi ⟩ \ ⟨F1 , . . . , Fi−1 ⟩ has a unique minimal element Ri .
2. For every i > 1, the complex Φi = ⟨Fi ⟩ ∩ ⟨F1 , . . . , Fi−1 ⟩ is pure of dimension d − 2.
14 12 34 24 23 13 15 25 35
1 3
2 1 4 2 3 5
5
∅
Figure 6.3 shows another example that shows how a shelling builds up a simplicial complex (in this case
the boundary of an octahedron) one step at a time. Note that each time a new triangle is attached, there is
a unique minimal new face.
Proposition 6.4.3. Let ∆d−1 be shellable, with h-vector (h0 , . . . , hd ). Then
hj = #{Fi : #Ri = j}
= # Fi : ⟨Fi ⟩ ∩ ⟨F1 , . . . , Fi−1 ⟩ has j faces of dimension d − 2 .
Moreover, if hj (∆) = 0 for some j, then hk (∆) = 0 for all k > j.
The proof is left as an exercise. One consequence is that the h-vector of a shellable complex is strictly non-
negative, since its coefficients count something. This statement is emphatically not true about the Hilbert se-
ries of arbitrary graded rings, or even arbitrary Stanley-Reisner rings of pure complexes (see Example 6.3.3
above).
149
a d a d
2 2
1 1
c b c b
a a d
3 4 3
2
1 1
c b c b e e
7
f a d f a d f a d
5 2 5 2 5 2
1 1 1
c b c b c b
4 3 6 4 3 6 4 3
e e e
8
i Fi Ri |Ri |
7 1 abc ∅ 0
f a d
2 abd d 1
5 2 3 bde e 1
1 4 bce ce 2
c b 5 acf f 1
6 3 6 cef ef 2
4
7 adf df 2
8 adf def 3
e
Figure 6.3: A step-by-step shelling of the octahedron with vertices a,b,c,d,e,f. Facets are labeled 1 . . . 8 in
shelling order. Enumerating the sets Ri by cardinality gives the h-vector (1, 3, 3, 1).
150
If a simplicial complex is shellable, then its Stanley-Reisner ring is Cohen-Macaulay (CM). This is an im-
portant and subtle algebraic condition that can be expressed algebraically in terms of depth or local co-
homology (topics beyond the scope of these notes) or in terms of simplicial homology (coming shortly).
Shellability is the most common combinatorial technique for proving that a ring is CM. The constraints on
the h-vectors of CM complexes are the same as those on shellable complexes, although it is an open problem
to give a general combinatorial interpretation of the h-vector of a CM complex.
Reisner’s theorem can be used to prove that shellable complexes are Cohen-Macaulay. The other ingredient
of this proof is a Mayer-Vietoris sequence, which is a standard tool in topology that functions sort of like
an inclusion/exclusion principle for homology groups, relating the homology groups of X, Y , X ∪ Y and
X ∩ Y . Here we can take X to be the subcomplex generated by the first n − 1 facets in shelling order and Y
the nth facet; the shelling condition says that the intersections and their links are extremely well-behaved,
so that Reisner’s condition can be established by induction on n.
Reisner’s theorem often functions as a working definition of the Cohen-Macaulay condition for combina-
torialists. The vanishing condition says that every link has the homology type of a wedge of spheres of
the appropriate dimension. (The wedge sum of a collection of spaces is obtained by identifying a point of
each; for example, the wedge of n circles looks like a flower with n petals. Reduced homology is additive
on wedge sums, so by (6.1) the wedge sum of n copies of Sd has reduced homology Rn in dimension d, and
0 in other dimensions.)
A Cohen-Macaulay complex ∆ is Gorenstein (over R) if in addition H̃dim ∆−dim σ−1 (link∆ (σ); R) ∼ = R for
all σ. That is, every link has the homology type of a sphere. This is very close to being a manifold. (I don’t
know offhand of a Gorenstein complex that is not a manifold, although I’m sure examples exist.)
A simplicial complex ∆ on vertex set E is a matroid complex if it is the family of independent sets of some
matroid M on E (see Defn. 3.4.1); in this case we write ∆ = I (M ). Many of the standard constructions of
matroid theory can be translated into simplicial complex language.
Say that a complex ∆ has property P hereditarily if every induced subcomplex ∆|X has property P; for
example, we have already seen that matroid complexes are hereditarily pure. (Note that the induced sub-
complex of ∆ on its entire vertex set is just itself, so if ∆ has P hereditarily then in particular it has P.)
Theorem 6.5.1. Let ∆ be an abstract simplicial complex on E. The following are equivalent:
151
2. ∆ is hereditarily shellable.
3. ∆ is hereditarily Cohen-Macaulay.
4. ∆ is hereditarily pure.
Proof. Work on this The implications (2) =⇒ (3) =⇒ (4) are consequences of the material in Chapter 6
(the first is a homework problem and the second is easy).
(4) =⇒ (1): Suppose I, J are independent sets with |I| < |J|. Then the induced subcomplex ∆|I∪J is pure,
which means that I is not a maximal face of it. Therefore there is some x ∈ (I ∪ J) \ I = J \ I such that
I ∪ x ∈ ∆, establishing (I3).
(1) =⇒ (4): Let F ⊆ E. If I is a non-maximum face of ∆|F , then we can pick J to be a maximum face, and
then (I3) says that there is some x ∈ J such that I + x is a face of ∆, hence of ∆|F .
To be written
6.7 Exercises
Problem 6.1. Let ∆ be a simplicial complex on vertex set V , and let v0 ̸∈ V . The cone over ∆ is the
simplicial complex C∆ generated by all faces σ + v0 for σ ∈ ∆.
Solution: (a) Every k-face of C∆ is either a k-face of ∆, or of the form σ + v0 where σ is a (k − 1)-face of ∆.
Therefore fk (C∆) = fk (∆) + fk−1 (∆) for all k. Equivalently, on the level of f -polynomials,
152
(expanding one factor of 1 − q in the first sum)
X X X
= fi−1 (∆)q i (1 − q)d+1 − fj−2 (∆)q j (1 − q)d−j+1 + fi−2 (∆)q i (1 − q)d+1−i
i j i
(c) Suppose F1 , . . . , Fn is a shelling order of ∆, i.e., each set ⟨F1 , . . . , Fi−1 ⟩ \ ⟨Fi ⟩ has a unique minimal face
Ri . Write σ̃ = σ + v0 for σ ∈ ∆; then
and then Ri is also the unique minimal face of the set on the right, so F̃1 , . . . , F̃n is a shelling order for C∆.
Conversely, suppose F̃1 , . . . , F̃n is a shelling order of ∆. If σ is any face of ⟨F̃1 , . . . , F̃i−1 ⟩ \ ⟨F̃i ⟩, then so is
σ \ {v0 }, so the unique minimal face Ri of that set does not contain v0 . Therefore
⟨F1 , . . . , Fi−1 ⟩ \ ⟨Fi ⟩ = ⟨F̃1 , . . . , F̃i−1 ⟩ \ ⟨F̃i ⟩ ∩ ∆
This argument provides a concrete explanation of why h(∆) = h(C∆), at least for shellable complexes.
Problem 6.2. Let ∆ be a graph (that is, a 1-dimensional simplicial complex) with c components, v vertices,
and e edges. Determine the isomorphism types of the simplicial homology groups H̃0 (∆; R) and H̃1 (∆; R)
for any coefficient ring R.
Solution: The simplest example I can think of is ∆ = ⟨12, 23, 34, 45⟩ and ∆ = ⟨12, 23, 13, 45⟩.
Problem 6.4. Prove that the two conditions in the definition of shellability (Defn. 6.4.1) are equivalent.
Solution: (1) =⇒ (2): Suppose that Ψi has a unique minimal element Ri (hence equals the interval
[Ri , Fi ]). Then
Φi = {σ ⊆ Fi : σ ̸⊇ Ri } = {σ ⊆ Fi : ∃v ∈ Ri : v ̸∈ σ}.
The maximal such faces σ are precisely those of the form Fi \ {w} for w ∈ Ri . These faces, which are all of
dimension d − 2, therefore generate Φi , so Φi is pure of dimension d − 2 as desired.
153
(2) =⇒ (1): Suppose that Φi is pure of dimension d − 2. Then every facet of Φi is a maximal proper
subface of Fi , i.e., Φi = ⟨Fi \ v1 , . . . , Fi \ vs ⟩ for some vertices v1 , . . . , vs ∈ Fi . Let R = {v1 , . . . , vs }; then
Ψi = {σ ⊆ Fi : σ ̸⊆ Fi \ vj ∀j ∈ [s]}
= {σ ⊆ Fi : vj ∈ σ ∀j ∈ [s]}
= {σ : R ⊆ σ ⊆ Fi }.
Solution: Since ∆d−1 is shellable, it decomposes into disjoint Boolean intervals [R1 , F1 ], . . . , [Rn , Fn ]. Let h̃i
be the number of intervals for which |Rj | = i. Therefore,
X X n
X X
fi−1 q i = q |σ| = q |σ|
i σ∈∆ j=1 σ∈[Rj ,Fj ]
n |Fj |
|Fj | − |Rj |
X X
= qk
j=1 k=|Rj |
k − |Rj |
d
d−i
X X
k
= h̃i q
i
k−i
k=i
d−i
ℓ d−i
X X
i
= h̃i q q (after substituting ℓ = k − i, k = ℓ + i)
i
ℓ
ℓ=0
X
= h̃i q i (1 + q)d−i .
i
or equivalently X X
fi q i = h̃i q i−1 (1 + q)d−i
i i
and comparing with equation (6.7) in the notes shows that h̃i = hi as desired.
Now, we want to show that the h-vector has no gaps. What I want to show is that if |Ri | = k > 0, then there
is some a < i for which |Ra | = k − 1. For i = 1 this is vacuous since |R1 | = 0, so suppose i ≥ 2.
Let Ri = {x1 , . . . , xk }. Then, for each j ∈ [k], the face Qj := Ri \ {xj } belongs to some interval [Raj , Faj ]
with aj < i. For convenience, reorder the vertices of Ri so that
a1 ≤ a2 ≤ · · · ≤ ak < i
and let a = ak . I claim that ak−1 < a. Indeed, if ak−1 = a, then Fa ⊃ Qk−1 ∩ Qk = Ri , which implies that
Ri ∈ ⟨F1 , . . . , Fa ⟩, a contradiction.
Let σ be any proper subface of Qk . Then σ is also a subface of some other Qj , so by the claim, it follows that
σ ∈ ⟨F1 , . . . , Fa−1 ⟩. But that is precisely the statement that Qk is a minimal element of ⟨Fa ⟩ \ ⟨F1 , . . . , Fa−1 ⟩,
so Qk = Ra , and of course |Qk | = k − 1 as desired.
154
Problem 6.6. Prove that the link operation commutes with union and intersection of complexes. That is, if
X, Y are simplicial complexes that are subcomplexes of a larger complex X ∪ Y , and σ ∈ X ∪ Y , then prove
that
linkX∪Y (σ) = linkX (σ) ∪ linkY (σ) and linkX∩Y (σ) = linkX (σ) ∩ linkY (σ).
Solution: First,
linkX∪Y (σ) = {τ ∈ X ∪ Y : τ ∩ σ = ∅, τ ∪ σ ∈ X ∪ Y }
= {τ ∈ X ∪ Y : τ ∩ σ = ∅, τ ∪ σ ∈ X} ∪ {τ ∈ X ∪ Y : τ ∩ σ = ∅, τ ∪ σ ∈ Y }
= {τ ∈ X : τ ∩ σ = ∅, τ ∪ σ ∈ X} ∪ {τ ∈ Y : τ ∩ σ = ∅, τ ∪ σ ∈ Y }
(because if τ ∪ σ ∈ Z then τ ∈ Z)
Problem 6.7. Let ∆ be a pure simplicial complex of dimension d − 1. ∆ is called shifted if its vertex set
can be labeled 1, . . . , n such that the following property holds: if σ ∈ ∆, j ∈ σ, i ∈
̸ σ, and i < j, then
σ \ {j} ∪ {i} ∈ ∆.
Equivalently, define a partial order ⪯ (called Gale order or componentwise order) on d-sets of positive integers
as follows: if a = (a1 < · · · < ad ) and b = (b1 < · · · < bd ), then a ⪯ b if ai ≤ bi for all i ∈ [d]. Then ∆ is
shifted if and only if its facets form an order ideal in Gale order.
Solution: (a) I claim that any linear extension σ1 , σ2 , . . . of Gale order on the facets is a shelling order.
Indeed, let σk = {a1 < · · · < ad } be a facet. We need to show that σk ∩ ⟨σ1 , . . . , σk−1 ⟩ is pure of dimension
d. Indeed, for any face τ ∈ σk ∩ ⟨σ1 , . . . , σk−1 ⟩, let aj = max(σk \ τ ) and ϕ = σ \ {aj }
If σk = [d] then k = 1 and there is nothing to prove. Otherwise, let j be the smallest vertex not in σk . I claim
that
σk ∩ ⟨σ1 , . . . , σk−1 ⟩ = ⟨σk \ {i} \ i > j⟩. (6.8)
To prove ⊇, observe that any such σk \ {i} is a face of σk as well as of σk \ {i} ∪ {j}, which precedes σk in
Gale order.
(b) A consequence of (6.8) is that σk ∩ ⟨σ1 , . . . , σk−1 ⟩ has exactly d + 1 − j facets, which implies that
hj (∆) = #{facets σ | min([n] \ σ) = d + 1 − j}.
155
(c) This is Theorem 2 of C. Klivans’s preprint “Shifted matroid complexes,” preprint, [Link]
Problem 6.8. (Requires some experience with homological algebra.) Prove that shellable simplicial com-
plexes are Cohen-Macaulay. (Hint: First do the previous problem. Then use a Mayer-Vietoris sequence.)
Solution: Suppose that ∆ is shellable. Every complex of dimension 0 is both shellable and Cohen-Macaulay,
and if dim ∆ = 1 then both shellability and CMness are equivalent to connectedness. So we can assume
dim ∆ ≥ 2.
Let F1 , . . . , Fn be a shelling order on the facets n. If n = 1 then ∆ is a simplex, which is trivially shellable
and Cohen-Macaulay. Otherwise, let Γ = ⟨F1 , . . . , Fn−1 ⟩ and F = Fn , so ∆ = Γ ∪ ⟨F ⟩, and by shellability,
the intersection Θ = Γ ∩ ⟨F ⟩ is a (d − 1)-dimensional complex generated by some of the maximal proper
faces of F .
Let σ be a face of dimension k. We want to show that link∆ σ (which has dimension d − k − 1) is APC. If
k = d then there is nothing to prove, so suppose k < d. By the Lemma we have a Mayer-Vietoris sequence
· · · → H̃i (linkΘ (σ)) → H̃i (linkΓ (σ)) ⊕ H̃i (link⟨F ⟩ (σ)) → H̃i (link∆ (σ)) → H̃i−1 (linkΘ (σ)) → · · ·
What are all these things? Certainly link⟨F ⟩ (σ) is either void (if σ ̸⊆ F ) or the simplex on F \ σ (otherwise);
in either case it is acyclic. So we can cross off all the Hi (link⟨F ⟩ (σ)) terms right away, obtaining
· · · → H̃i (linkΘ (σ)) → H̃i (linkΓ (σ)) → H̃i (link∆ (σ)) → H̃i−1 (linkΘ (σ)) → · · · (6.9)
Let F = {v0 , . . . , vd }, so that (after suitable relabeling) Θ = Γ ∩ ⟨F ⟩ = ⟨F \ v0 , . . . , F \ vr ⟩ for some r ∈ [0, d].
So (
∅ if σ ̸⊆ F,
linkΘ (σ) =
⟨F \ vi \ σ : 0 ≤ i ≤ r, vi ̸∈ σ⟩ if σ ⊆ F.
In the second case, this complex is generated by some (possibly empty) subset of the maximal proper facets
of the (d − k − 1)-simplex F \ σ. Therefore, its homology is either zero, or concentrated in degree d − k − 2.
Therefore, the term H̃i (link∆ (σ)) in (6.9) is stuck between zeros for all i < d − k − 1, hence is zero. We have
just shown that the homology of link∆ (σ) is concentrated in dimension d − k − 1, as desired.
Problem 6.9. Complete the proof of Theorem 6.5.1 by showing that hereditarily pure simplicial complexes
are shellable. (Hint: Pick a vertex v. Show that the two complexes
∆1 = del∆ (v) = ⟨σ ∈ ∆ : v ̸∈ σ⟩,
∆2 = link∆ (v) = ⟨σ − v ∈ ∆ : v ∈ σ⟩
are both shellable. Then concatenate the shelling orders to produce a shelling order on ∆. You will probably
need Problem 6.1.) As a consequence of the construction, derive a relationship among the h-polynomials of
∆, ∆1 , and ∆2 .
Solution: Induct on dimension and on n = |E|. For the base case, every complex with one vertex is trivial,
as is (more generally) every complex of dimension 0. If all vertices are cone points then ∆ is a simplex and
there is nothing to prove, so suppose that v is a vertex that is not a cone point. Since ∆1 = ∆|E−v is an
induced subcomplex of ∆, it inherits local purity, and it has fewer vertices, hence is shellable by induction.
Now we show that ∆2 is hereditarily pure as well, hence also shellable by induction:
156
• If dim ∆|A = dim ∆|A+v then every facet of ∆|A is a facet of ∆|A+v . In this case, the facets of ∆2 |A are
the faces α − v, where α is a facet of ∆|A+v containing v. Hence ∆2 |A is pure, since ∆|A+v is pure.
• If dim ∆|A = dim ∆|A+v − 1 then ∆|A+v is the cone over ∆|A by v. So ∆2 |A = ∆|A is again pure.
By the way, given a total order on the vertices v1 < · · · < vn , we can iterate this process: first split up at
vn , then at vn−1 , etc. The resulting shelling order on ∆ is called the reverse lexicographic order5 : F < F ′ if
max(F △F ′ ) ∈ F ′ . For example, the reverse lexicographic order on U2 (5) is
12, 13, 23, 14, 24, 34, 15, 25, 35, 45.
(Despite the appearance of homology, all you really need is the rank-nullity theorem from linear algebra.
The choice of ground field k is immaterial, but you can take it to be R if you want.)
Solution: Suppose dim ∆ = d − 1. Consider the simplicial chain complex of ∆ (everything in sight is a
k-vector space, and I will drop k from the notation):
k ∂k+1 ∂
0 → Cd−1 → Cd−2 → · · · → Ck+1 −−−→ Ck −→ Ck−1 → · · · → C0 → C−1 → 0.
Then
X
χ̃(∆) = (−1)k dim Ck (∆)
k
X
= (−1)k (rank ∂k + nullity ∂k ) (by the rank-nullity theorem)
k
X
= (−1)k (dim im ∂k + dim ker ∂k )
k
X
= (−1)k (dim ker ∂k − dim im ∂k+1 ) (reindexing)
k
X
= (−1)k dim H̃k (∆).
k
5 This terminology is standard, but is confusing since it is not in fact the reverse of lexicographic order. To avoid this problem, some
157
Problem 6.11. Express the h-vector of a matroid complex in terms of the Tutte polynomial of the underlying
matroid. (Hint: First figure out a deletion/contraction recurrence for the h-vector, using Problem 6.9.)
Solution: For a matroid M on ground set E, write h(M ) = h(M, q) = h0 +h1 q+h2 q 2 +· · · , where (h0 , h1 , . . . )
is the h-vector of its independence complex ∆ = I (M ) (see Chapter 6). We claim that for any e ∈ E,
1 if E = ∅,
h(M \e) if e is a loop,
h(M ) = (6.12)
h(M/e) if e is a coloop,
h(M \e) + qh(M/e) otherwise.
The first case is trivial. Deleting a loop does not change the independence complex, and contracting a coloop
corresponds to decoding, which does not change the h-vector by Problem 6.1. For the general case, recall
from Problem 6.9 that a shelling order for ∆ is given as follows: first list the facets of del∆ (e) = I (M \e) in
shelling order, then list the facets of link∆ (e) = I (M/e) in shelling order. If the decompositions associated
with the shellings of the deletion and link were
G G
del∆ (e) = [Ri , Fi ], link∆ (e) = [Si , Gj ]
i j
from which the fourth case of (6.12) follows by Proposition 6.4.3. Now, the Tutte Recipe Theorem for Ma-
troids (Theorem 4.2.1) gives
h(M, q) = TM (1, q).
Problem 6.12. Let V = {x11 , x12 , . . . , xn 1, xn2 }. Consider the simplicial complex
∆n = {σ ⊆ V : σ ̸⊆ {xi , yi } ∀i ∈ [n]}.
(In fact, ∆n is the boundary sphere of the crosspolytope, the convex hull of the standard basis vectors and
their negatives in Rn .) Determine the f - and h-polynomials of ∆n .
∆(c1 , . . . , cn ) = {σ ⊆ V : |σ ∩ Vi | ≤ 1 ∀i ∈ [n]}.
(The previous problem is the case that ci = 2 for all i.) Show that ∆(c1 , . . . , cn ) is shellable. Determine its
f - and h-polynomials.
Solution: The complete colorful complex X = ∆(c1 , . . . , cn ) is shellable because it is a matroid complex:
specifically, the independence complex of the matroid M = U1 (c1 ) × · · · × Un (cn ). So we could, if we like,
apply Problem 6.11:
n n n
X Y Y Y 1 − q cj
h(X, q) = hi (X)q i = TM (1, q) = TUj (cj ) (1, q) = (1 + q + q 2 + · · · + q cj −1 ) = .
i j=1 j=1 j=1
1−q
158
Meanwhile, using (6.7),
X X
fi (X)q i = hi (X)q i−1 (1 + q)d−i
i i
i
(1 + q)d X
q
= hi (X)
q i
1+q
c j
q
(1 + q)d Y 1 − 1+q
n
= q
q j=1
1 − 1+q
n cj
(1 + q)d+1
Y q
= 1−
q j=1
1+q
The i-faces of ∆n are the simplices of the form {±ei0 , . . . , ±eii } where 1 ≤ i0 < i1 < · · · < ii ≤ n. Therefore,
n
fi (Xn ) = 2i+1
i+1
where the binomial coefficient counts the choices for {i0 , . . . , ii } and the power of 2 counts the choices for
signs.
159
Chapter 7
Polytopes include familiar objects such as cubes, pyramids, and Platonic solids. They are central in linear
programming and therefore in optimization, and exhibit a wealth of nice combinatorics. The classic book
on polytopes is Grünbaum [Grü03]; an equally valuable, more recent reference is Ziegler [Zie95]. A good
reference for the basics is chapter 2 of Schrijver’s notes [Sch13].
First some key terms. A subset S ⊆ Rn is convex if, for any two points in S, the line segment joining them
is also a subset of S. The smallest convex set containing a given set T is called its convex hull, denoted
conv(T ). Explicitly, one can show (Problem 7.2; not hard) that
r
( )
X
conv(x1 , . . . , xr ) = c1 x1 + · · · + cr xr : 0 ≤ ci ≤ 1 for all i and ci = 1 . (7.1)
i=1
These points are called convex linear combinations of the xi . A related definition is the affine hull of a
point set, which is the smallest affine linear space containing it:
r
( )
X
aff(x1 , . . . , xr ) = c1 x1 + · · · + cr xr : ci = 1 . (7.2)
i=1
The interior of S as a subspace of its affine span is called the relative interior of S, denoted relint S. This
concept is necessary to talk about interiors of different-dimensional polyhedra in a sensible way. For ex-
ample, the closed line segment S = {(x, 0) : 0 ≤ x ≤ 1} in R2 has empty interior as a subset of R2 , but its
affine span is the x-axis, so relint S = {(x, 0) : 0 < x < 1}.
Clearly conv(T ) ⊆ aff(T ) (in fact, the inclusion is strict if 1 < |T | < ∞). For example, the convex hull of
three non-collinear points in Rn is a triangle, while their affine hull is the unique plane (i.e., affine 2-space)
containing that triangle.
Definition 7.1.1. A polyhedron P is a nonempty intersection of finitely many closed half-spaces in Rn .
Equivalently,
P = {x ∈ Rn : ai1 x1 + · · · + ain xn ≥ bi ∀i ∈ [m]}
where aij , bi ∈ R. These equations are often written as a single matrix equation Ax ≥ b, where A ∈ Rm×n
and b ∈ Rm .
160
Definition 7.1.2. A polytope is a bounded polyhedron.
The “Fundamental Theorem of Polytopes” asserts that polytopes are precisely the sets in Rn that can be
expressed as the convex hull of a finite set of points. We will prove this theorem in the next section.
Definition 7.1.3. A point v in a polyhedron P is a vertex of P if v ̸∈ conv(P \ {v}). The set of all vertices of
P will be denoted by V (P ).
Definition 7.1.4. Let P ⊆ Rn be a polyhedron. A face of P is a subset F ⊆ P that maximizes some linear
functional ℓ : Rn → R, i.e., ℓ(x) ≥ ℓ(y) for all x ∈ F , y ∈ P . In this case, we write F = maxP (ℓ). The face is
proper if ℓ is not a constant. The dimension of a face is the dimension of its affine span.
The only improper face is P itself. Note that the union of all proper faces is the topological boundary ∂P
(proof left as an exercise).
To make this a bit more concrete, suppose P is a polytope in R3 . What point or set of points is highest? In
other words, what points maximize the linear functional (x, y, z) 7→ z? The answer to this question might
be a single vertex, or an edge, or a polygonal face. Of course, there is nothing special about the z-direction.
For any direction given by a linear functional ℓ, the extreme points of P in that direction are by definition
the maxima of the linear functional x 7→ ℓ(x), and the set of those points forms a face of P .
For a linear functional chosen “at random”, the face it determines will almost surely be a vertex of P .
Higher-dimensional faces correspond to more special directions.
Proposition 7.1.5. Let P ⊆ Rn be a polyhedron. Then:
Proof. (1) Each face F is defined by adding a linear inequality to the list of inequalities defining P . Specifi-
cally, if F = maxP (ℓ), and ℓ(x) = m for all x ∈ F , then F = {x ∈ P : ℓ(x) ≥ m}.
(2) Let F ′ = maxP (ℓ′ ) and F ′′ = maxP (ℓ′′ ) and suppose that F ′ ∩ F ′′ contains a point x. Let F = maxP (ℓ),
where ℓ = ℓ′ + ℓ′′ (in fact any positive linear combination of ℓ′ , ℓ′′ will do). Then x is a global maximum of ℓ
on P , and since x also maximizes both ℓ′ and ℓ′′ , the face F consists exactly of those points of P maximizing
both ℓ′ and ℓ′′ . In other words, F = F ′ ∩ F ′′ , as desired.
(3) By (2), the desired face Fx is the intersection of all faces containing x.
(4) If x ∈ ∂Fx then Fx has a face G containing x, but G is also a face of P by (1), which contradicts the
definition of Fx .
P
(5) Suppose that x is a 0-dimensional
P face, i.e., {x} = maxP (ℓ). If x is a convex linear combination ci y i
of points yi ∈ P , then ℓ(x) ≥ ci ℓ(yi ), with equality only if ℓ(yi ) = ℓ(x) for all i. But then yi = x for all i
by assumption. Therefore x ̸∈ conv(P \ {x}), hence is a vertex.
161
On the other hand, if x ∈ P is not an 0-dimensional face, then by (4) x ∈ relint Fx . Then Fx contains a ball
centered at x, hence a line segment centered at x, and thus x is a linear combination of the two endpoints
of the segment (namely, their average). Hence x is not a vertex.
(6) The set F (P ) is certainly a bounded poset under inclusion, and by (2) it is a meet-semilattice, hence a
lattice by Prop. 1.2.9.
To prove that it is ranked, we need a new construction. Let v be a vertex maximized by some linear func-
tional ℓ, so that we can find a constant c such that
P/v
H
P
One can show [Zie95, Prop. 2.4] that P/v is a polytope of dimension dim(P ) − 1, and that there is a bijection
regardless of the particular choice of ℓ and H. In particular, the face lattice F (P/v) is isomorphic to the
interval [v, P ] ⊂ F (P ).
P ∗ := {y ∈ Rn | x · y ≤ 1 ∀x ∈ P }. (7.3)
Observe that P ∗ is bounded. (P contains a ball of radius ϵ centered at the origin, which implies that P ∗ is
contained in a ball of radius 1/ϵ.) Moreover, by definition P ∗ is the intersection of half-spaces, but it is not
clear at this point that it is the intersection of finitely many of them — we will prove that in the next section
(and give an example then).
Temporarily, say that a subset of Rn is a P-polytope if it is the convex hull of a finite set of points.
162
Theorem 7.2.1 (The Fundamental Theorem of Polytopes). A set P ⊆ Rn is a polytope (i.e., a bounded polyhe-
dron) if and only if it is a P-polytope.
The proof occupies this whole section, and will include several proofs of claims along the way.
First, let P be the intersection of finitely many half-spaces, i.e., P = {x ∈ Rn : Ax ≤ b}, where A ∈ Rm×n
and b ∈ Rm×1 . By projecting onto the orthogonal complement of the rowspace of A, we can assume WLOG
that rank A = n. For each point x ∈ P , let Ax be the submatrix of A consisting of rows ai for which ai ·x = bi .
(These rows correspond to linear functionals maximized at x.)
Claim 1: Let x ∈ P . Then {x} = maxP (ℓ) for some ℓ ∈ (Rn )∗ if and only if rank Ax = n. (More generally,
rank Ax = n − dim Fx .)
n ∗
P 1. If rank Ax = n then there is a basis {λ1 , . . . , λn } for (R ) such that x ∈ maxP (λi ) for each i.
Proof of Claim
Let λ = ci λi , where c1 , . . . , cn > 0. Since the λi form a basis, if x is any other point in P , then there is
some i such that λi (x) ̸= λi (x); by assumption this must mean λi (x) < λi (x), and so λ(x) < λ(x). It follows
that maxP (λ) = {x}.
Now suppose rank Ax < n. For convenience, reorder the rows of A so that a1 , . . . , ar are the rows of Ax
and ar+1 , . . . , am are the remaining rows; in particular
aj · x < bj ∀j ∈ [r + 1, m]. (7.4)
The system of equations {ai · x = bi : i ∈ [r]} defines an affine space of dimension n − rank Ax > 0.
Let v be any vector parallel to that affine space (equivalently, perpendicular to each of a1 , . . . , am ). Then
ai ·(x+ϵv) = bi for any ϵ ∈ R. By continuity and (7.4), we can choose ϵ > 0 small enough that aj ·(x±ϵv) < bj
for all j > r. Then x′ = x + ϵv and x′′ = x + ϵv belong to P , and for any linear functional ℓ, either
ℓ(x′ ) ≤ ℓ(x) ≤ ℓ(x′′ ) or ℓ(x′′ ) ≤ ℓ(x) ≤ ℓ(x′ )
(depending on the sign of ℓ(v)), so that x cannot be the unique maximum of ℓ on P .
In general, let x be a point that is not a vertex, and let Fx be the unique minimal face of P containing x. By
assertion (3) of Prop. 1.26, x is in the relative interior of Fx , so it is a convex combination
ℓ
X ℓ
X
x= ci yi , 0 ≤ ci ≤ 1, ci = 1 (7.5)
i=1 i=1
163
for each i ∈ [ℓ]. Plugging (7.6) into (7.5) gives
ℓ k k ℓ
!
X X X X
x = ci bij vj = ci bij vj . (7.7)
i=1 j=1 j=1 i=1
Note that
ℓ
X ℓ
X
0≤ ci bij ≤ ci ≤ 1
i=1 i=1
so formula (7.7) is an expression for x as a convex combination of the v. (Summary of calculation: A convex
combination of convex combinations is a convex combination.)
At this point, we have shown that every polytope is a P-polytope. Part 2 of the proof is to show the converse.
Let P ⊂ Rn be a polytope. By Part 1 of the proof, we know that P has finitely many vertices v1 , . . . , vr and
we can write P = conv(v1 , . . . , vr ). Assume without loss of generality that aff(P ) = Rn (otherwise, replace
Rn with the affine hull) and that the origin is in the interior of P (translating if necessary), so that we can
consider the polar dual P ∗ = {y ∈ Rn | x · y ≤ 1 ∀x ∈ P } (see Definition 7.1.6).
Claim 3:
P ∗ = {y ∈ Rn | vi · y ≤ 1 ∀i ∈ [r]}.
r
X r
X
x= ci vi , 0 ≤ ci ≤ 1, ci = 1,
i=1 i=1
whence
r
X r
X
x·y = ci vi · y ≤ ci = 1 ∴ y ∈ P ∗.
i=1 i=1
This claim enables us to draw pictures of duals. For example, consider the polytope
x ≥ −2
P = (x, y) ∈ R2 | y ≥ −1 = conv {(−2, 2), (−2, −1), (2, −1)} .
3x + 4y ≤ 2
From the V-description of P and Claim 3, we can easily read off the H-description of the dual:
−2x + 2y ≤ 1
P ∗ = (x, y) ∈ R2 | −2x − y ≤ 1 .
2x − y ≤ 1
INSERT FIGURE
164
In particular, P ∗ is an intersection of finitely many half-spaces. So, by the first part of the theorem, P ∗ is a
P-polytope, say P ∗ = conv{y1 , . . . , ys }. Meanwhile, the double dual P ∗∗ = (P ∗ )∗ is defined by
P ∗∗ = {x ∈ Rn : x · y ≤ 1 ∀y ∈ P ∗ }
(7.8)
= {x ∈ Rn : x · yj ≤ 1 ∀j ∈ [s]}
where the second equality comes from Claim 3.
Claim 4: P = P ∗∗ .
Pr
Proof of Claim 4. First, we show that P ⊆ P ∗∗ . Let x ∈ P = conv(z1 , . . . , zr ), say x = i=1 ci zi , and let
j ∈ [s]. Then
Xr r
X
x · yj = ci zi · yj ≤ ci = 1
i=1 i=1
∗ ∗∗
since zi · yj ≤ 1 for all i, j by definition of P . Therefore x ∈ P .
(where the < comes from (7.9) and the ≤ from the hypothesis x ∈ P ∗∗ ), a contradiction.
Consequently, (7.8) expresses P as the intersection of finitely many half-spaces, and we have shown that
the P-polytope P is in fact a polytope.
• A facet of P is a face of codimension 1 (that is, dimension n − 1). In this case there is a unique linear
functional (up to scaling) that is maximized on F , given by the outward normal vector from P . Faces
of codimension 2 are called ridges and faces of codimension 3 are sometimes called peaks.
1 This seemingly obvious assertion is not so easy to prove, although it is true in more generality: if S ⊆ Rn is a convex set and
y ̸∈ S, then there exists a hyperplane separating S from y — or equivalently a linear functional ℓ : Rn → R such that ℓ(y) > 0
and ℓ(x) < 0 for all x ∈ S. This is called Minkowski’s Hyperplane Separation Theorem. It is equivalent to many other statements,
including Farkas’ Lemma.
165
• A supporting hyperplane of P is a hyperplane that meets P in a nonempty face.
• P is simplicial if every face is a simplex. For example, every 2-dimensional polytope is simplicial, but
of the Platonic solids in R3 , only the tetrahedron, octahedron and icosahedron are simplicial — the
cube and dodecahedron are not. The boundary of a simplicial polytope is thus a simplicial (n − 1)-
sphere.
• P is simple if every vertex belongs to exactly n faces. (In fact no vertex can belong to fewer than n
faces.)
Proposition 7.3.2. 1. F (P ∗ ) = F (P )∗ (i.e., the dual of F (P ) in the sense of Definition 1.1.12).
2. A polytope P is simple if and only if its dual P ∗ is simplicial.
3. A polytope is both simple and simplicial if and only if it is a simplex.
Proof. To be written.
One of the big questions about polytopes is to classify their possible f -vectors and, more generally, the
structure of their face posets. Here is a result of paramount importance.
Theorem 7.4.1. Let ∆ be the boundary sphere of a convex simplicial polytope P ⊆ Rd . Then ∆ is shellable, and its
h-vector is a palindrome, i.e., hi = hd−i for all i.
These equations are the Dehn-Sommerville relations. They were first proved early in the 20th century, but
the following proof, due to Bruggesser and Mani [BM71], is undoubtedly the one in the Book.
Sketch of proof. Let F1 , . . . , Fn be the facets of P and let Ai = aff(Fi ). Let ℓ be a line that passes through the
interior of P and meets the hyperplanes Ai in n distinct points. (Note that almost any line will do.) Imagine
walking along ℓ, starting inside P . When you get to infinity, Stage 1 ends and Stage 2 starts by “hopping”
to the other side of the line and come back the other way until you get back to inside P . Relabel the facets in
the order that you encounter their affine spans. Let Am be the last affine span you cross before “hopping”.
As you keep walking, keep your eyes on P . In Stage 1, after crossing A1 , all you can see is F1 , but for each
i ∈ {2, . . . , m}, the facet Fi pops into view as soon as you cross Ai . (You have probably already seen some
of its boundary, but not the entire facet.) In Stage 2, facets Am+1 , . . . , An are visible, but after you cross Ai
the facet Fi disappears from view (although some of its boundary may still be visible). Finally, just before
you cross An and enter P again, all you can see is Fn .
F3 F3
F4
F1 F1
F5
F2 F2
Stage 1 Stage 2
166
In fact F1 , . . . , Fn is a shelling order (called a line shelling), because
(
⟨ridges of Fj that are visible before crossing Aj ⟩ for 2 ≤ j ≤ m,
⟨Fj ⟩ ∩ ⟨F1 , . . . , Fj−1 ⟩ =
⟨ridges of Fj that are invisible after crossing Aj ⟩ for m + 1 ≤ j ≤ n
(This assertion does need to be checked.) A shelling of P coming from a line in this way is called a line
shelling. Moreover, since every ridge belongs to exactly two facets, we observe that each facet Fi con-
tributes to hk (P ), where
On the other hand, the reversal of this shelling order is also a line shelling (by traversing ℓ in the opposite
direction). Since each facet shares a ridge with exactly d other facets (because P is simplicial!), the previous
formula says that if a facet contributes to hi with respect to the original shelling order, then it contributes to
hd−i in the reverse shelling order. The h-vector is an invariant of P , so it follows that hi = hd−i for all i.
The Dehn-Sommerville relations are a basic tool in classifying h-vectors, and therefore f -vectors, of simpli-
cial polytopes. Since h0 = 1 for shellable complexes, it follows immediately that the only possible h-vectors
for simplicial polytopes in R2 and R3 are (1, k, 1) and (1, k, k, 1), respectively (where k is a positive integer),
and in particular the number of facets determines the h-vector (which is not the case in higher dimensions).
Recall from Definition 7.1.4 that a face of a polyhedron P ⊂ Rn is defined as the subset of P that maximizes
a linear functional. We can get a lot of mileage out of classifying linear functionals by which face of P they
maximize. The resulting structure N (P ) is called the normal fan of P . (Technical note: officially N (P ) is a
structure on the dual space (Rn )∗ , but we typically identify (Rn )∗ with Rn by declaring the standard basis
to be orthonormal — equivalently, letting each vector in Rn act by the standard dot product.)
Given a face F ⊂ P , let σF be the collection of linear functionals maximized on F . As we will see, the sets
σF are in fact the interiors of cones (convex unions of rays from the origin).
Example 7.5.1. Let P = conv{(1, 1), (1, −1), (−1, 1)} ⊂ R2 . The polytope and its normal fan are shown
below.
The word “fan” means “collection of cones”. Multiplying a linear functional by a positive scalar does not
change the face on which it is maximized, and that if ℓ and ℓ′ are linear functionals maximized on the same
face, then so is every functional aℓ + bℓ′ , where a, b are positive scalars. Therefore, each σF is a cone. The
vertices x, y, z correspond to the 2-dimensional cones, the edges Q, R, S to 1-dimensional cones (a.k.a. rays)
and the polytope P itself to the trivial cone consisting of the origin alone. In general, if F is a face of a
polytope P ⊆ Rn , then
dim σF = n − dim F. (7.10)
◀
167
σR
x R z
σz
Q σx
σQ
S
σy
y σS
P N (P )
Example 7.5.2 (The normal fan of an unbounded polyhedron). Let P be the unbounded polyhedron defined
by the inequalities x ≤ 1, y ≤ 1, x + y ≤ 1 (so its vertices are x = (0, 1) and y = (1, 0)). The polytope and its
normal fan are shown below.
Q y
σQ
R
σR
x
σy
σx
S σS
P N (P )
This normal fan is incomplete: it does not cover every linear functional in (R2 )∗ , only the ones that have a
well-defined maximum on P (in this case, those in the first quadrant). It is not hard to see that the normal
fan of a polyhedron is complete if and only if the polyhedron is bounded, i.e., a polytope. The dimension
formula for normal cones (7.10) is still valid in the unbounded case. ◀
In general the normal fan of a polytope can be quite complicated, and there exist fans in Rn that are not the
normal fans of any polytope, even for n = 3; see, e.g., [Zie95, Example 7.5]. However, for some polytopes,
we can describe the normal fan using other combinatorics, such as the following important class.
168
The theory of generalized permutahedra is usually considered to have started with Postnikov’s paper [Pos09];
other important sources include [PRW08] and [AA17]. Edmonds [Edm70] considered equivalent objects
earlier under the name “polymatroids” (add details).
Theorem 7.5.4. A polytope P ⊆ Rn is a generalized permutahedron if and only if every edge of P is parallel to
ei − ej for some i, j, where {e1 , . . . , en } is the standard basis.
Generalized permutahedra can also be described as certain degenerations of the standard permutahedron,
which is the convex hull of the vectors (w1 , . . . , wn ), where w ranges over all permutations of [n]. The
normal fan of the standard permutahedron is precisely the braid fan.
One important family of generalized permutahedra are matroid base polytopes. Given a matroid M on
ground set [n], let P be the convex hull of all characteristic vectors of bases of M . It turns out that P is a
generalized permutahedron; in fact, the matroid base polytopes are exactly the generalized permutahedra
whose vertices are 0/1 vectors [GGMS87, Thm. 4.1]. Describing the faces of matroid polytopes in terms of
the combinatorics of the matroid is an interesting and difficult problem; see [FS05].
The central problem considered in this section is the following: How many integer or rational points are
in a convex polytope? An excellent and comprehensive source is [BR15]. There is some material on the
case of a rational polytope in [Sta12, §4.6.2].
Definition 7.6.1. A polytope P ⊆ RN is integral (resp. rational) if and only if all vertices of P have integer
(resp. rational) coordinates.
For a set P ⊆ RN and a positive integer n, let nP = {nx : x ∈ P }. (nP is called a dilation of P .)
The (relative) boundary of P , written ∂P , is the union of proper faces of P , that is, the set of points x ∈ P
such that for every ε > 0, the ball of radius ε (its intersection with aff(P )) contains both points of P and
points not in P .
i(P, n) = |nP ∩ ZN |
i∗ (P, n) = |n(relint P ) ∩ ZN |
i(P,
a n)a is the number of integer points in nP or, equivalently, the number of rational points in P of the form
0 1 aN
, ,..., . Our goal is to understand the functions i(P, n) and i∗ (P, n).
n n n
We start with P a simplex, and with an easy example. Let
Then
nP = conv{(0, 0, 0), (n, n, 0), (n, 0, n), (0, n, n)}.
169
Case 1. If the αi are all integers, the resulting points are integer points and the sum of the coordinates is
even. How many such points are there? The answer is the number of monomials in four variables of degree
n, that is, n+3
3 . However, there are other integer points in nP .
Case 2. We can allow the fractional part of αi to be 1/2. If any one of the αi has fractional part 1/2, the others
must be also. Writing γi = αi − 1/2, we get points of the form
(γ1 + 1/2)(1, 1, 0) + (γ2 + 1/2)(1, 0, 1) + (γ3 + 1/2)(0, 1, 1)
= γ1 (1, 1, 0) + γ2 (1, 0, 1) + γ3 (0, 1, 1) + (1, 1, 1).
P P P
Note here that γi = ( αi ) − 3/2 ≤ n − 3/2. Since the γi are integers, γi ≤ n − 2. Sothe number of
these points equals the number of monomials in four variables of degree n − 2, that is, n+1 3 .
Note that all the points in Case 2 are interior points because each αi = γi + 1/2 > 0 and their sumP is at most
n − 2 + 3/2 (less than n). A point in Case 1 is an Pinterior point if and only if all the αi > 0 and αi < n.
The four-tuples (α1 − 1, α 2 − 1, α3 − 1, n − 1 − α i ) correspond to monomials in four variables of degre
n − 4; there are n−1
3 of them. Thus we get
n−1
n+1 1 5
i∗ (P, n) = + = n3 − n2 + n − 1,
3 3 3 3
another polynomial. (Anything else you notice? Is it a coincidence?)
170
Thus Q is a half-open parallelepiped containing 0 and P̃ .
Proposition 7.6.3. Let P be an integral N -simplex in RN , with vertices v0 , v1 , . . . , vN , and let C = C(P̃ ). A
PN
point z ∈ ZN +1 is an integer point in C if any only if z = y + i=0 ri (v1 , 1) for some y ∈ Q ∩ ZN +1 and some
nonnegative integers ri . Furthermore, this representation of z is unique.
So to count integer points in C (and hence to determine i(P, n)), we only need to know how many integer
points are in Q with each fixed (integer) last coordinate. We call the last coordinate of z ∈ Q the degree of z.
PN
Note that for z ∈ Q, deg z = i=0 ri for some ri , 0 ≤ ri < 1, so if deg z is an integer, 0 ≤ deg z ≤ N .
Theorem 7.6.4. Let P be an integral N -simplex in RN , with vertices v0 , v1 , . . . , vN , let C = C(P̃ ), and let
PN
Q = { i=0 ri (vi , 1) : for each i, 0 ≤ ri < 1}. Let δj be the number of points of degree j in Q ∩ ZN +1 . Then
∞
X δ0 + δ1 λ + · · · + δN λN
i(P, n)λn = .
n=0
(1 − λ)N +1
Proof.
∞
X
i(P, n)λn = (δ0 + δ1 λ + · · · + δN λN )(1 + λ + λ2 + · · · )N +1
n=0
∞ !
N
X k+N k
= (δ0 + δ1 λ + · · · + δN λ ) λ .
N
k−0
PN n−j+N
The coefficient of λn on the right hand side is
j=0 δj N .
For the interior of P (and of C) we use an analogous construction, but with the opposite half-open paral-
lelipiped. Let
(N )
X
Q∗ = ri (vi , 1) : 0 < ri ≤ 1∀i .
i=0
Proposition 7.6.6. Let P be an integral N -simplex in RN , with vertices v0 , v1 , . . . , vN , and let C = C(P̃ ). A point
PN
z ∈ ZN +1 is an integer point in relint C if and only if z = y + i=0 ci (v1 , 1) for some y ∈ Q∗ ∩ ZN +1 and some
nonnegative integers ci . Furthermore, this representation of z is unique.
So to count integer points in relint C (and hence to determine i∗ (P, n)), we only need to know how many
integer points are in Q∗ with each fixed (integer) last coordinate. Note that for z ∈ Q∗ , 0 < deg z ≤ N + 1.
Theorem 7.6.7. Let P be an integral N -simplex in RN , with vertices v0 , v1 , . . . , vN , let C = C(P̃ ), and let
PN
Q∗ = { i=0 ri (vi , 1) : for each i, 0 < ri ≤ 1}. Let δj∗ be the number of points of degree j in Q∗ ∩ ZN +1 . Then
∞
X δ1∗ λ + δ2∗ λ2 + · · · + δN
∗
+1 λ
N +1
i∗ (P, n)λn = .
n=0
(1 − λ)N +1
171
Now the punchline is that there is an easy relationship between the δi and the δi∗ . Note that
(N )
X
∗
Q = ri (vi , 1) : for each i, 0 < ri ≤ 1
i=0
N
( )
X
= (1 − ti )(vi , 1) : for each i, 0 ≤ ti < 1
i=0
N N
( )
X X
= (vi , 1) − ti (vi , 1) : for each i, 0 ≤ ti < 1
i=0 i=0
N N
!
X X
= (vi , 1) − Q = vi , N + 1 −Q
i=0 i=0
So far I have considered only integral simplices. To extend the result to integral polytopes requires trian-
gulation of the polytope, that is, subdivision of the polytope into simplices. The extension is nontrivial. We
cannot just add up the functions i and i∗ for the simplices in the triangulation, since interior points of the
polytope can be contained in the boundary of a simplex of the triangulation, and in fact in the boundary of
more than one simplex of the triangulation. But it works in the end.
Theorem 7.6.10. Let P ⊆ RN be an integral polytope of dimension N . Then
∞
X
(1 − λ)N +1 i(P, n)λn
i=0
PN j
As before, write this polynomial as j=0 δj λ . What can we say about the coefficients δj ?
δ0 = i(P, 0) = 1, since this is the number of integer points in the polytope 0P = {0}.
172
I claim C is the volume of P . To see this, note that vol(nP ) = nN vol(P ) (if P is of full dimension N ). Now
the volume of nP can be estimated by the number of lattice points in nP , that is, by i(P, n). In fact,
i(P, n)
So C = lim = vol(P ).
n→∞ nN
One last comment. The Ehrhart theory can be generalized to rational polytopes. In the more general case,
the functions i(P, n) and i∗ (P, n) need not be polynomials, but are quasipolynomials—restricted to a con-
gruence class in some modulus (depending on the denominators occurring in the coordinates of the ver-
tices) they are polynomials. An equivalent description is that the function i(P, n) is a polynomial in n and
expressions of the form gcd(n, k), e.g.,
(
(n + 1)2 n even
i(P, n) = = (n + gcd(n, 2) − 1)2 .
n2 n odd
7.7 Exercises
Problem 7.1. Prove that the topological boundary of a polyhedron is the union of its proper faces.
Solution: If p is a point on a proper face maximized by some nonconstant linear functional ℓF , then any
open ball centered at p contains some points with greater values of ℓ, hence not in P . On the other hand, if
p is not on any proper face, then ℓF (p) < ℓF (F ) for all faces F ; there are finitely many faces so there is some
ϵ for which ℓF (p) + ϵ < ℓF (F ) for all F , and it follows that the ball of radius ϵ/2 centered at p is in P ◦ .
Problem 7.2. Prove that the convex hull of a finite point set X = {x1 , . . . , xn } ⊆ Rd is the set of convex
linear combinations of it.
P
Solution: Suppose p is a convex linear combination of the xi , say p = ci xi . We prove that p ∈ conv(X)
by induction on r(p) = #{i ∈ [n] : ci ̸= 0}, using the fact that every line segment between two points in
X is contained in conv(X). If r(p) = 1, then p = xi for some i and we are done. Otherwise, suppose that
ci , cj > 0, where i ̸= j; then let
X X
qi = (ci + cj )xi + ck xk , qj = (ci + cj )xj + ck xk
k̸=i,j k̸=i,j
For theP P that the set C of all convex linear combinations of X is itself a convex set. Indeed,
converse, we claim
if p = ci xi and q = di xi are points in C, then for any λ ∈ [0, 1] we have
n
X
λp + (1 − λ)q = (λci + (1 − λ)di )xi
i=1
173
P P
and the coefficients (λci + (1 − λ)di ) are positive and add up to λ ci + (q − λ) di = λ + (1 − λ) = 1.
Therefore the line segment between p and q is contained in C. On the other hand, C is a convex set
containing X, so by definition it contains conv(X).
Problem 7.3. Prove Theorem 7.5.4.
Solution: Suppose the normal fan NP coarsens the braid arrangement. Let ℓ(x) = c · x; if c1 , . . . , cn are
distinct then ℓ lies in the interior of a maximal-dimensional normal cone, hence is maximized at a unique
vertex. Conversely, if F = maxP (ℓ) > 0 is a face of dimension > 0, then ci = cj for some i, j. In particular,
walking along F in the direction ei − ej keeps ℓ constant, so if F is an edge it must be parallel to ei − ej .
UNDER CONSTRUCTION. Conversely, if every edge is parallel to ei − ej for some i, j, then no linear
functional with distinct coordinates is constant on any edge. That is, every maximal cone in the normal fan
(i.e., every normal cone of a vertex) is a union of braid cones. (But what about non-maximal cones?)
Let F be a face of positive dimension, let ℓ ∈ NP (F ) be given by ℓ(x) = c · x. For every pair i, j such that
some edge of F is parallel to ei − ej , the fact that ℓ is constant on F implies that ci = cj . We want to show
that every other functional in the same braid cone as ℓ is also maximized on F .. . .
Problem 7.4. Let M be a matroid and let P be its base polytope. Prove that P is a generalized permutahe-
dron in two different ways:
(a) Show that the normal fan NP coarsens the braid cone, using the properties of greedy algorithms.
(b) Show that every edge of P is parallel to some ei − ej , using the properties of basis exchange.
Solution: (1) Let ℓ(x) = c · x be a linear functional. The face of M maximized by ℓ is the convex hull of the
vertices of M maximized by ℓ, which are the possible outputs when the greedy algorithm (§3.4) is run using
c1 , . . . , cn as the weights. but this algorithm never does any arithmetic involving the ci ’s; it only compares
them. Therefore the possible output depends only on the order of the numbers ci , i.e., the face of the braid
arrangement containing c.
(2) Suppose that F is an edge of P whose endpoints correspond to two distinct bases B and B ′ . Then there
is a linear functional ℓ(x) = c · x such that ℓ(B) = ℓ(B ′ ) > ℓ(B ′′ ) for all other bases B ′′ . Let i be an element
of B△B ′ for which ci is as small as possible; without loss of generality i ∈ B \ B ′ . By basis exchange, there
exists j ∈ B ′ \ B such that B ′′ = B − i + j is a basis. But ℓ(B − i + j) = ℓ(B) − ci + cj ≥ ℓ(B) by the choice of
i; by maximality B ′′ must be either B or B ′ , and it certainly cannot be B so it must be B ′ . We have shown
that B ′ = B − i + j, and the edge F is therefore parallel to the vector ei − ej .
Problem 7.5. Let ∆n be the n-dimensional simplex whose vertices are the n standard basis vectors in Rn ,
together with the origin. That is,
∆n = {x = (x1 , . . . , xn ) ∈ Rn : 0 ≤ xi ≤ 1, 0 ≤ x1 + · · · + xn ≤ 1}.
Calculate the Ehrhart polynomials i(∆n , k) and i∗ (∆n , k).
arrangements of k stars and n bars, in which xi is the number of consecutive stars before the ith bar. For
example, if k = 4 and n = 3, then ∗| ∗ ∗ ∗ | ∗ | corresponds to the point (1, 3, 1), while ∗ ∗ | ∗ ∗||∗ corresponds
to (2, 2, 0). (Note that stars after the last bar do not count.) Therefore,
(k + n)(k + n − 1) · · · (k + 2)(k + 1)
k+n
i(∆, k) = = .
n n(n − 1) · · · (2)(1)
174
Meanwhile, the interior of the kth dilate is
and so
= |{(y1 , . . . , yn ) ∈ Zn : 0 ≤ yi ≤ k − n − 1, 0 ≤ y1 + · · · + yn ≤ k − n − 1}|
k−1
=
n
(by the same stars-and-bars argument; now there are k − n − 1 stars and n bars)
(k − 1)(k − 2) · · · (k − n + 1)(k − n)
= .
n(n − 1) · · · (2)(1)
Solution:
(because the S summand depends only on s = |S|, so we may replace each S with [s])
n
X n
= 2s |{(x1 , . . . , xs ) ∈ Zs : 1 ≤ xi ≤ k, 0 ≤ x1 + · · · + xs ≤ k}|
s=0
s
n
X n
= 2s |{(x1 , . . . , xs ) ∈ Zs : 1 ≤ xi ≤ k − s + 1, s ≤ x1 + · · · + xs ≤ k}|
s=0
s
175
(removing impossible cases)
n
X n
= 2s |{(y1 , . . . , ys ) ∈ Zs : 0 ≤ yi ≤ k − s, 0 ≤ y1 + · · · + ys ≤ k − s}|
s=0
s
(setting yi = xi − 1)
n
X n k
= 2s
s=0
s s
n
n k(k − 1) · · · (k − s + 1)
X
= 2s .
s=0
s s!
Meanwhile
n
n k−1
X
∗ s
i (Xn , k) = i(Xn , k − 1) = 2 .
s=0
s s
176
Chapter 8
Group Representations
Some remarks:
• ρ specifies an action of G on V that respects its vector space structure. So we have all the accou-
trements of group actions, such as orbits and stabilizers. If there is only one representation under
consideration, it is often convenient to use group-action notation and write gv (or g · v, g(v), etc.)
instead of the bulkier ρ(g)v.
• It is common to say that ρ is a representation, or that V is a representation, or that the pair (ρ, V ) is a
representation.
• ρ is faithful if it is injective as a group homomorphism.
Example 8.1.2. Let G be any group. The trivial representation is the map
ρtriv : G → GL1 (k) ∼
= k×
sending g 7→ 1 for all g ∈ G. (This is as non-faithful as you can get.) ◀
Example 8.1.3. Let kG be the vector space of formal k-linear combinations of elements of G: that is, kG =
h∈G h h : ah ∈ k . The regular representation of G is the map ρreg : G → GL(kG) defined by
P
a
!
X X
g ah h = ah (gh).
h∈G h∈G
That is, g permutes the standard basis vectors of kG according to the group multiplication law. Thus
dim ρreg = |G|. This representation is faithful. ◀
177
The vector space kG is a ring, with multiplication given by multiplication in G and extended k-linearly. In
this context it is called the group algebra of G over k.
Remark 8.1.4. A representation of G is equivalent to a (left) module over the group algebra kG. Technically
“representation” refers to the way G acts and “module” refers to the space on which it acts, but the two
terms really carry the same information.
Example 8.1.5. Let G = Sn , the symmetric group on n elements. The defining representation ρdef of G
on kn maps each permutation σ ∈ G to the n × n permutation matrix with 1’s in the positions (i, σ(i)) for
every i ∈ [n], and 0’s elsewhere. This representation is faithful. ◀
Example 8.1.6. More generally, let G act on a finite set X. Then there is an associated permutation repre-
sentation on kX, the vector space with basis X, given by
!
X X
g ax x = ax (gx).
x∈X x∈X
For short, we might specify the action of G on X and say that it “extends linearly” to kX. For instance, the
action of G on itself by left multiplication gives rise in this way to the regular representation, and the usual
action of Sn on an n-element set gives rise to the defining representation. ◀
Example 8.1.7. Let G = Z/nZ be the cyclic group of order n, and let ζ ∈ k be a nth root of unity (not
necessarily primitive). Then G has a 1-dimensional representation given by ρ(x) = ζ x . This representation
is faithful if and only ζ is a primitive root of unity. In a sense, every representation of G is built from these
1-dimensional representations, as we will see later (§8.8). ◀
Example 8.1.8. Consider the dihedral group Dn of order 2n, i.e., the group of symmetries of a regular n-gon,
given in terms of generators and relations by
extended by group multiplication (e.g., ρgeo (sr2 ) = ρgeo (s)ρgeo (r)2 , etc.).
2. We can regard Dn as the symmetries of a regular n-gon with vertices labeled 1, 2, . . . , n in cyclic order.
We can take r to be the rotation by 2π/n, We can take s to be the reflection across the line L through
the center and vertex n. Note that if n is even then L also passes through vertex n/2, while if n is odd
then L passes through the midpoint of the side with vertices ⌊n/2⌋, ⌈n/2⌋. Thus the action is given by
the homomorphism Dn → Sn defined by
178
Example 8.1.9. The symmetric group Sn has another 1-dimensional representation, the sign representa-
tion, given by (
1 if σ is even,
ρsign (σ) =
−1 if σ is odd.
The sign representation is nontrivial provided char k ̸= 2. Note that ρsign (g) = det ρdef (g) (see Exam-
ple 8.1.5). (More generally, if ρ is any representation, then det ρ is a 1-dimensional representation.) ◀
Example 8.1.10. Let (ρ, V ) and (ρ′ , V ′ ) be representations of G. The direct sum ρ ⊕ ρ′ : G → GL(V ⊕ V ′ ) is
defined by
(ρ ⊕ ρ′ )(g)(v + v ′ ) = ρ(g)(v) + ρ′ (g)(v ′ )
for v ∈ V , v ′ ∈ V ′ . In terms of matrices, (ρ ⊕ ρ′ )(g) is a block-diagonal matrix:
ρ(g) 0
.
0 ρ′ (g)
For example, both V and V ′ are G-invariant subspaces of V ⊕ V ′ . For a more subtle example, let G = Sn
and let (ρdef , kn ) be the defining representation. The one-dimensional subspace spanned by the vector
(1, 1, . . . , 1) is fixed point wise by Sn , so it is a G-invariant subspace that carries the trivial representation.
When are two representations the same? More generally, what is a map between representations?
Definition 8.2.1. Let (ρ, V ) and (ρ′ , V ′ ) be representations of G. A linear transformation ϕ : V → V ′ is
G-equivariant, or for short a G-map, if ρ′ (g) · ϕ(v) = ϕ(ρ(g) · v) for all g ∈ G and v ∈ V . (more concisely,
gϕ = ϕg for all g ∈ G). An equivalent condition is that the following diagram commutes:
ϕ
V V′
ρ(g) ρ′ (g) (8.1)
ϕ
V V′
179
We sometimes use the notation ϕ : ρ → ρ′ for a G-equivariant map. In the language of modules, a G-
equivariant transformation is the same thing as a homomorphism of G-modules.1
is G-equivariant because permuting the coordinates of a vector does not change their sum. ◀
Example 8.2.4. Let n be odd, and consider the dihedral group Dn acting on a regular n-gon (see Exam-
ple 8.1.8). Label the vertices 1, . . . , n in cyclic order. Label each edge the same as its opposite vertex, as in
the figure on the left. Then the permutation action ρV on vertices is isomorphic to the action ρE on edges.
In other words, the diagram on the right commutes for all g ∈ Dn , where “opp” is the map that sends the
basis vector for a vertex to the basis vector for its opposite edge.
5 4
opp
kn kn
3 1
ρV (g) ρE (g)
1 3
kn kn
opp
4 2 5
The case that n is even is trickier, because then each reflection either fixes two vertices or two edges, but not
both. ◀
Example 8.2.5. Let v1 , . . . , vn be the points of a regular n-gon in R2 centered at the origin, e.g., vj =
cos 2πj 2πj
n , sin n . Then the map R → R sending the jth standard basis vector to vj is Dn -equivariant,
n 2
180
so that ρV ◦ ϕ(g) = ϕ ◦ ρE (g) for all g, i.e., the following diagram commutes:
ϕ
E V
ρE (g) ρV (g)
ϕ
E V
◀
Kernels and images of G-equivariant maps are well-behaved. Those familiar with modules will not be
surprised: every kernel or image of a R-module homomorphism is also a R-module. In category-theoretic
terms, the family of R-modules forms an abelian category.
Proposition 8.2.7. Any G-equivariant map ϕ : (ρ, V ) → (ρ′ , V ′ ) has G-invariant kernel and G-invariant image.
On the other hand, the representation σ = ρtriv ⊕ ρsign on V (see Examples 8.1.2 and 8.1.9) is given by
1 0 1 0
σ(id) = , σ(flip) = .
0 1 0 −1
These two representations ρ and σ are in fact isomorphic. Indeed, ρ acts trivially on k⟨e1 + e2 ⟩ and acts
by the sign representation on k⟨e1 − e2 ⟩. These two vectors form a basis of V (here is where we use the
assumption char k ̸= 2), and one can check that the change-of-basis map
−1 " 1 1
#
1 1 2 2
ϕ= = 1
1 −1 − 21
2
If ρ = ρ′ , then any linear transformation ϕ : k → k (i.e., any map ϕ(v) = cv for some c ∈ k) will do. Thus
the set of G-equivariant homomorphisms is actually isomorphic to k.
Assume char k ̸= 2 (otherwise ρtriv = ρsign and we are done at this point). If ϕ : ρtriv → ρsign is G-
equivariant, then we have diagrams
ϕ ϕ
V V′ V V′
ρtriv (12) ρsign (12) ρtriv (21) ρsign (21)
ϕ ϕ
V V′ V V′
181
The first diagram always commutes because ρtriv (12) = ρsign (12) is the identity map, but the second dia-
gram says that for every v ∈ k
and since char k ̸= 2 we are forced to conclude that c = 0. Therefore, there is no nontrivial G-homomorphism
ρtriv → ρsign . ◀
Example 8.2.9 is the tip of an iceberg: we can use the vector space HomG (ρ, ρ′ ) of G-homomorphisms
ϕ : ρ → ρ′ to measure how similar ρ and ρ′ are.
Clearly, every representation can be written as the direct sum of indecomposable representations, and every
irreducible representation is indecomposable. On the other hand, there exist indecomposable representa-
tions that are not irreducible.
Example 8.3.2. As in Example 8.2.8, let V = {e1 , e2 } be the standard basis for k2 , where char k ̸= 2. Recall
that the defining representation of S2 = {id, flip} is given by
1 0 0 1
ρdef (id) = , ρdef (flip) =
0 1 1 0
Fortunately, we can rule out this kind of pathology most of the time.
Theorem 8.3.3 (Maschke’s Theorem). Let G be a finite group, let k be a field whose characteristic does not di-
vide |G|, and let (ρ, V ) be a representation of G over k. Then every G-invariant subspace has a G-invariant comple-
ment. In particular, (ρ, V ) is semisimple.
Proof. If ρ is irreducible, then there is nothing to prove. Otherwise, let W be a nontrivial G-invariant sub-
space, and let π : V → W be a projection, i.e., a linear surjection that fixes the elements of W pointwise.
182
(Such a map π can be constructed as follows: choose a basis for W , extend it to a basis for V , and let π fix
all the basis elements in W and kill all the ones in V \ W .)
The map π is k-linear, but not necessarily G-equivariant. However, we can turn π into a G-equivariant
projection by “averaging over G”. (This trick will come up again and again.) Define π̃ : V → W by
1 X
π̃(v) = gπ(g −1 v). (8.2)
|G|
g∈G
1 X
π̃(hv) = gπ(g −1 hv)
|G|
g∈G
1 X
= (hk)π((hk)−1 hv)
|G|
k∈G: hk=g
1 X
= h kπ(k −1 v) = hπ̃(v).
|G|
k∈G
Maschke’s Theorem implies that, when the conditions hold, a representation ρ is determined up to iso-
morphism by the multiplicity of each irreducible representation in ρ (i.e., the number of isomorphic copies
appearing as direct summands of ρ). Accordingly, to understand representations of G, we should first study
irreps.
Example 8.3.4. Let k have characteristic 0 (for simplicity), and G = Sn . The defining representation of G
on kn (Example 8.1.5) is not simple. The one-dimensional subspace L = ⟨(1, 1, . . . , 1)⟩ is fixed pointwise by
every permutation σ ∈ Sn , and is therefore an invariant subspace, carrying the trivial representation.2 By
Maschke’s theorem, L has a G-invariant complement. In fact, L⊥ is the orthogonal complement of L under
the standard inner product on kn , namely the space of all vectors whose coordinates sum to 0. This is called
(a little confusingly) the standard representation of Sn , denoted ρstd . That is,
Thus dim ρstd = n − 1. We will soon be able to prove that ρstd is irreducible (Problem 8.2). ◀
2 For the same reason, every permutation representation of every group has a trivial summand.
183
8.4 Characters
The first miracle of representation theory is that we can detect the isomorphism type of a representation ρ
without knowing every matrix ρ(g): it turns out that we just need to know their traces.
Definition 8.4.1. Let (ρ, V ) be a representation of G over k. Its character is the function χρ : G → k given
by
χρ (g) = tr ρ(g).
Note that characters are in general not group homomorphisms.
Example 8.4.2. Some simple facts and some characters we’ve seen before:
• A one-dimensional representation is its own character. (In fact these are exactly the characters that
are homomorphisms.)
• For any representation ρ, we have χρ (IdG ) = dim ρ, because ρ(IdG ) is the n × n identity matrix.
• The defining representation ρdef of Sn has character
χdef (σ) = number of fixed points of σ.
Indeed, this is true for any permutation representation of any group.
• The regular representation ρreg has character
(
|G| if σ = IdG
χreg (σ) = ◀
0 otherwise.
Example 8.4.3. Consider the geometric representation ρgeo of the dihedral group Dn = ⟨r, s : rn = s2 =
e, srs = r−1 ⟩ by rotations and reflections:
1 0 cos θ sin θ
ρgeo (s) = , ρgeo (r) = .
0 −1 − sin θ cos θ
The character of ρgeo is
Proposition 8.4.4. Characters are class functions; that is, they are constant on conjugacy classes of G. Moreover,
if ρ ∼
= ρ′ , then χρ = χρ′ .
Now, let ϕ : ρ → ρ′ be an isomorphism represented by an invertible matrix Φ; then Φρ(g) = ρ′ (g)Φ for all
g ∈ G and so Φρ(g)Φ−1 = ρ′ (g). Taking traces gives χρ = χρ′ .
184
Surprisingly, the converse of the second assertion is also true: a representation is determined up to isomor-
phism by its character!
The basic vector space functors3 of direct sum, duality, tensor product and Hom carry over naturally to
representations, and behave well on their characters. Throughout this section, let G be a finite group and
let (ρ, V ) and (ρ′ , W ) be finite-dimensional representations of G over C, with V ∩ W = 0.
1. Direct sum. To construct a basis for V ⊕ W , we can take the union of a basis for V and a basis for W .
Equivalently, we can write the vectors in V ⊕ W as column block vectors:
v
V ⊕W = : v ∈ V, w ∈ W .
w
(hϕ)(v) = ϕ(h−1 v)
Proof. Let J be the Jordan canonical form of ρ(h) (which exists since we are working over C), so that χρ(h) =
tr J. The diagonal entries Jii are its eigenvalues, which must be roots of unity since h has finite order, so
their inverses are their complex conjugates. Meanwhile, J −1 is an upper-triangular matrix with (J −1 )ii =
(Jii )−1 = Jii , and tr J −1 = χρ∗ (h) = χρ (h).
3. Tensor product. Fix bases {e1 , . . . , en } and {f1 , . . . , fm } for V and W respectively. As a vector space,
define4
V ⊗ W = k ⟨ei ⊗ fj | 1 ≤ i ≤ n, 1 ≤ j ≤ m⟩ ,
3 Technically, a functor is an operation that not only takes objects to objects, but also maps to maps. E.g., applying duality to vector
spaces not only takes V and W to V ∗ and W ∗ , but also takes every linear transformation ϕ : V → W to a linear transformation
ϕ∗ : W ∗ → V ∗ . We will not be too concerned with functors as such in this section, but will mostly focus on their effect on characters/
4 The “official” definition of the tensor product is much more functorial and can be made basis-free, but this concrete definition is
185
equipped with a multilinear action of k (that is, c(x ⊗ y) = cx ⊗ y = x ⊗ cy for c ∈ k). In particular,
dim(V ⊗ W ) = (dim V )(dim W ). We can accordingly define a representation (ρ ⊗ ρ′ , V ⊗ W ) by
or more concisely
h · (v ⊗ w) = hv ⊗ hw
extended bilinearly to all of V ⊗ W .
That is, the left-hand side is the entry in the row corresponding to ei ⊗ fj and column corresponding to
ek ⊗ fℓ . In particular,
n
! m
X X X
χρ⊗ρ′ (h) = (ρ(h))i,i (ρ′ (h))j,j = (ρ(h))i,i (ρ′ (h))j,j = χρ (h)χρ′ (h). (8.4)
(i,j)∈[n]×[m] i=1 j=1
First, the vector space HomC (V, W ) admits a representation of G, in which h ∈ G acts on a linear transfor-
mation ϕ : V → W by sending it to the map h · ϕ defined by
(h · ϕ)(v) = h(ϕ(h−1 v)) = ρ′ (h) ϕ(ρ(h−1 )(v)) . (8.5)
for h ∈ G, ϕ ∈ HomC (V, W ), v ∈ V . It is straightforward to verify that this is a genuine group action, i.e.,
that (hh′ ) · ϕ = h · (h′ · ϕ).
Moreover, HomC (V, W ) ∼ = V ∗ ⊗ W as vector spaces and G-modules. To see this, suppose that dim V = n
and dim W = m; then the elements of V ∗ and W can be regarded as 1 × n and m × 1 matrices respectively
(the former acting on V , which consists of n × 1 matrices, by matrix multiplication). Then the previous
description of tensor product implies that V ∗ ⊗ W consists of m × n matrices, which correspond to elements
of HomC (V, W ). This isomorphism is G-equivariant by (8.5), so
so we may consider these functors as operations on characters: e.g., (χ ⊗ ψ)(g) = χ(g)ψ(g), etc.
What about HomG (V, W )? Evidently HomG (V, W ) ⊆ HomC (V, W ), but equality need not hold. For ex-
ample, if V and W are the trivial and sign representations of Sn (for n ≥ 2), then HomC (V, W ) ∼
= C but
HomG (V, W ) = 0. (See Example 8.2.9.)
186
The two Homs are related as follows. In general, when a group G acts on a vector space V , the subspace of
G-invariants is defined as
V G = {v ∈ V | hv = v ∀h ∈ G}.
This is the largest subspace of V that carries the trivial action.
Observe that a linear map ϕ : V → W is G-equivariant if and only if hϕ = ϕ for all h ∈ G, where G acts on
HomC (V, W ) as in (8.5). (The proof of this fact is left to the reader.)
Moreover, G acts by the identity on HomG (ρ, ρ′ ), so its character is a constant function whose value is
dimC HomG (ρ, ρ′ ). To put it another way, the action of G is a sum of copies of the trivial representation. We
want to understand how many.
From now on, we assume that k = C (though everything would be true over an algebraically closed field
of characteristic 0), unless otherwise specified.
Recall that a class function is a function χ : G → C that is constant on conjugacy classes of G. The set Cℓ(G)
of class functions forms a vector space. Define an inner product on Cℓ(G) by
1 X 1 X
⟨χ, ψ⟩G = χ(h) ψ(h) = |C| χ(C) ψ(C) (8.8)
|G| |G|
h∈G C
where C runs over all conjugacy classes. Observe that ⟨·, ·⟩G is a sesquilinear form (i.e., C-linear in the
second term and conjugate linear in the first). It is also non-degenerate, because the indicator functions of
conjugacy classes form an orthogonal basis for Cℓ(G). Analysts might want to regard the inner product as
a convolution (with summation over G as a discrete analogue of integration).
1 X
dimC V G = χρ (h) = χtriv , χρ G
.
|G|
h∈G
Proof. The second equality follows from the definition of the inner product. For the first equality, define a
linear map π : V → V by
1 X
π= ρ(h).
|G|
h∈G
G
Note that π(v) ∈ V for all v ∈ V , because
1 X 1 X
gπ(v) = ghv = ghv = π(v)
|G| |G|
h∈G gh∈G
187
That is, π is a projection from V → V G . Choose a basis for V consisting of a basis for V G and extend it to a
basis for V . With respect to that basis, π can be represented by the block matrix
I ∗
0 0
1 X
dimC V G = tr(π) = χρ (h).
|G|
h∈G
By the way, we know by Maschke’s Theorem that V is semisimple, so we can decompose it as a direct sum
of irreps. Then V G is precisely the direct sum of the irreducible summands on which G acts trivially.
Example 8.6.2. Let G act on a set X, and let ρPbe the corresponding permutation representation onP the space
CX. For each orbit O ⊆ X, the vector vO = x∈O x is fixed by G. On the other hand, any vector x∈X ax x
fixed by G must have ax constant on each orbits. Therefore the vectors vO are a basis for V G , and dim V G is
the number of orbits. So Proposition 8.6.1 becomes
1 X
# orbits = # fixed points of h
|G|
h∈G
1 X 1 X
χρ (h)χρ′ (h) = χHom(ρ,ρ′ ) (h) by (8.6)
|G| |G|
h∈G h∈G
One intriguing observation is that this expression is symmetric in ρ and ρ′ , since ⟨α, β⟩G = ⟨β, α⟩G in
general, but dimC HomG (ρ, ρ′ ) is real. (It is not algebraically obvious that HomG (ρ, ρ′ ) and HomG (ρ′ , ρ)
should have equal dimension.)
Proposition 8.6.4 (Schur’s Lemma). Let G be a group, and let (ρ, V ) and (ρ′ , V ′ ) be finite-dimensional irreps of G
over a field k (not necessarily of characteristic 0).
That is, the only G-equivariant maps from an G-irrep to itself are multiplication by a scalar.
188
Proof. For (1), recall from Proposition 8.2.7 that ker ϕ and im ϕ are G-invariant subspaces. But since ρ, ρ′
are simple, there are not many possibilities. Either ker ϕ = 0 and im ϕ = W , when ϕ is an isomorphism.
Otherwise, ker ϕ = V or im ϕ = 0, either of which implies that ϕ = 0.
We can now prove the following omnibus theorem, which essentially reduces the study of representations
of finite groups to the study of characters.
Theorem 8.6.5 (Fundamental Theorem of Character Theory for Finite Groups). Let (ρ, V ) and (ρ′ , V ′ ) be
finite-dimensional representations of G over C.
then
n
X
χρ , χρi G
= mi and χρ , χρ G
= m2i .
i=1
and consequently
n
X
(dim ρi )2 = |G|. (8.10)
i=1
5. The irreducible characters (i.e., characters of irreps) form an orthonormal basis for Cℓ(G). In particular, the
number of irreducible characters equals the number of conjugacy classes of G.
Proof. For assertion (1), the equation (8.9) follows from part (2) of Schur’s Lemma together with Proposi-
tion 8.6.3. It follows that the characters of isomorphism classes of irreps are an orthonormal basis for some
subspace of the finite-dimensional space Cℓ(G), so there can be only finitely many of them. (This result
continues to amaze me every time I think about it.)
Assertion (2) follows because the inner product is additive on direct sums. That is, every class function ψ
satisfies
⟨χρ⊕ρ′ , ψ⟩G = ⟨χρ + χρ′ , ψ⟩G = ⟨χρ , ψ⟩G + ⟨χρ′ , ψ⟩G .
189
For (3), Maschke’s Theorem says that every complex representation ρ can be written as a direct sum of
irreducibles. Their multiplicities determine ρ up to isomorphism, and can be recovered from χρ by (2).
(Again, amazing.)
For (4), recall that χreg (IdG ) = |G| and χreg (g) = 0 for g ̸= IdG . Therefore
1 X 1
χreg , ρi G
= χreg (g)ρi (g) = |G|ρi (IdG ) = dim ρi
|G| |G|
g∈G
For (5) Schur’s Lemma together with assertion (3) imply that the irreducible characters are orthonormal
in Cℓ(G), hence linearly independent. The trickier part is to show that they in fact span Cℓ(G). Let Y be
the subspace of Cℓ(G) spanned by the irreducible characters, and let
n o
Z = Y ⊥ = ϕ ∈ Cℓ(G) : ϕ, χρ G = 0 for every irreducible character ρ .
or equivalently
1 X
Tρ (v) = ϕ(g)gv
|G|
g∈G
(to parse this, note that ϕ(g) is a number). Our plan is to show that Tρ is the zero map (in disguise), then
deduce that ϕ = 0.
Second, we show that Tρ = 0. We start with the case that ρ is irreducible. Since Tρ is G-equivariant, Schur’s
Lemma implies that it is multiplication by a scalar. On the other hand,
1 X
tr(Tρ ) = ϕ(g)χρ (g) = ϕ, χρ G
= 0
|G|
g∈G
because ϕ ∈ Z = Y ⊥ . Since Tρ has trace zero and is multiplication by a scalar, it is the zero map. (Again,
here we need k to have characteristic 0.)
190
Now we consider an arbitrary representation ρ. By Maschke’s Theorem, every complex representation ρ is
semisimple, and the definition of Tρ implies that it is additive on direct sums (that is, Tρ⊕ρ′ = Tρ + Tρ′ ),
proving that Tρ = 0 for all ρ.
1 X
0 = Tρreg (IdG ) = ϕ(g)ρreg (g).
|G|
g∈G
This is an equation in the vector space of |G| × |G| matrices. Observe that the permutation matrices ρreg (g)
have disjoint supports (because the only group element that maps h to k is kh−1 ), hence are linearly inde-
pendent. Therefore ϕ(g) = 0 for all g, so ϕ is the zero map.
We have now shown that Y has trivial orthogonal complement as a subspace of Cℓ(G), so Y = Cℓ(G),
completing the proof.
Ln Ln
be the irreps of G, and let α = i=1 ρ⊕a
Corollary 8.6.6. Let ρ1 , . . . , ρn P i
i
and β = i=1 ρ⊕b
i
i
be two representa-
n
tions. Then dim HomG (α, β) = i=1 ai bi , and in particular HomG (α, β) is a direct sum of constant maps between
irreducible summands of α and β.
The irreducible characters of G can thus be written as a square matrix X with columns corresponding to
conjugacy classes. This is called the character table of G. (It is helpful to include the size of the conjugacy
class in the table, for ease of computing scalar products.) By orthonormality of characters, the matrix X is
close to being unitary: XDX ∗ = I, where the star denotes conjugate transpose and D is the diagonal matrix
with entries |C|/|G|. It follows that X ∗ X = D−1 , which is equivalent to the following statement:
Proposition 8.6.7. Let χ1 , . . . , χn be the irreducible characters of G and let C, C ′ be any two conjugacy classes.
Then
n
X |G|
χi (C)χi (C ′ ) = δC,C ′ .
i=1
|C|
This is a convenient tool because it says that different columns of the character table are orthogonal under
the usual scalar product (without having to correct for the size of conjugacy classes).
At this point, we can also take care of some unfinished business: we claimed in Example 8.3.4 that the
standard representation ρstd of Sn (defined by ρdef = ρstd ⊕ ρtriv ) is irreducible for all n ≥ 2. This claim can
now be proved using characters. We leave the proof as an exercise (Problem 8.2), but will use the fact freely
in what follows. Note that the standard character maps each permutation to its number of fixed points
minus one.
Theorem 8.6.5 provides the basic tools to calculate character tables. In general, the character table of a
finite group G with k conjugacy classes is a k × k table in which rows correspond to irreducible char-
acters χ1 . . . . , χk and columns to conjugacy classes. Part (1) of the Theorem says that the rows form an
orthonormal basis under the inner product on class functions, so computing a character table resembles a
Gram-Schmidt process. The hard part is coming up with enough representations whose characters span
Cℓ(G). Here are some ways of generating them:
191
• Every group carries the trivial and regular characters, which are easy to write down. The regular
character contains at least one copy of every irreducible character.
• The symmetric group also has the sign and standard characters, both of which are irreducible.
• Many groups come with natural permutation actions whose characters can be added to the mix.
• The operations of duality and tensor product can be used to come up with new characters. Duality
preserves irreducibility, but tensor product typically does not.
Remark 8.7.1. While we have defined Cℓ(G) as a vector space, we are really concerned with the Z-module
CℓZ (G) generated by the irreducible characters. Its elements are called virtual characters (so called because
a virtual character is a sum of irreducible characters with some negative coefficients, then it does not come
from a genuine representation). Being a basis for CℓZ (G) as a Z-module (an integral basis) is a much
stronger condition than being a basis for Cℓ(G) as a vector space.
In the following examples, we will notate a character χ by a bracketed list of its values on conjugacy classes,
in the same order that they are listed in the table. Numerical subscripts will always be reserved for irre-
ducible characters.
Example 8.7.2. The group G = S3 has three conjugacy classes, indexed by cycle shapes:
We know three irreducible characters of S3 : the trivial and sign characters, which are both 1-dimensional,
and the standard character, which is 2-dimensional. So this must be the complete character table:
1 3 2
C111 C21 C3
χtriv 1 1 1
χsign 1 −1 1
χstd 2 0 −1
Note that the regular character χreg = [6, 0, 0] equals χtriv + χsign + 2χstd , as predicted by Theorem 8.6.5(e).
It is also worth noting that χsign ⊗ χstd = χstd . ◀
Example 8.7.3. We calculate all the irreducible characters of S4 . There are five conjugacy classes, corre-
sponding to the cycle-shapes 1111, 211, 22, 31, and 4. The squares of their dimensions must add up to
|S4 | = 24; the only list of five positive integers with that property is 1, 1, 2, 3, 3.
We do know three irreducible characters, namely χtriv , χsign , and χstd , of dimensions 1,1,3 respectively. So
we need to come up with two more. The regular character is always useful, and we can also try out the
character χstd ⊗ χsign :
192
1 6 3 8 6
C1111 C211 C22 C31 C4
χ1 = χtriv 1 1 1 1 1
χ2 = χsign 1 −1 1 1 −1
χ3 = χstd 3 1 −1 0 −1
χ4 = χstd ⊗ χsign 3 −1 −1 0 1
χreg 24 0 0 0 0
It is easy to check that ⟨χ4 , χ4 ⟩G = 1, so χ4 is irreducible. A general fact illustrated here is that tensoring
with the sign character preserves irreducibility — for an even more general statement, see Problem 8.1.
The other irreducible character χ5 has dimension 2. We can calculate it from the regular character and the
other four irreducibles, because
and so
χreg − χ1 − χ2 − 3χ3 − 3χ4
χ5 = .
2
Therefore the complete character table of S4 is as follows.
1 6 3 8 6
C1111 C211 C22 C31 C4
χ1 1 1 1 1 1
χ2 1 −1 1 1 −1
χ3 3 1 −1 0 −1
χ4 3 −1 −1 0 1
χ5 2 0 2 −1 0
◀
Example 8.7.4. For a non-symmetric example, let G be the dihedral group D4 = ⟨r, s : r4 = s2 = e, srs =
r−1 ⟩ of order 8. The elements e, r, r2 , r3 are rotations (of orders 1,4,2,4 respectively) and the elements
s, sr, sr2 , sr3 are reflections (all of order 2).
First we determine the conjugacy classes. The identity is of course its own conjugacy class, and r2 is central
(it commutes with everything) so it is also its own conjugacy class. The two order-4 rotations are conjugate
by the defining relation of D4 . That leaves the reflections. On the one hand, rsr−1 = rsr3 = (rsr)r2 = sr2 ,
which says that s is conjugate to sr2 , and a similar calculation shows that sr and sr3 are conjugate. On
the other hand, conjugation by s preserves the parity of the number of r’s in a word, so these are the only
conjugacies. In summary, there are five conjugacy classes:
Therefore there are 5 irreps, and the squares of their dimensions must sum to 8; the only possibility is
1,1,1,1,2.
Now let’s collect some characters. There’s always the trivial character and the regular character. By the
parity argument, sending r 7→ 1 and s 7→ −1 is a homomorphism G → C× ; we’ll call this the orientation
character. From Example 8.1.8, we also have the characters of the geometric representation ρgeo and the
193
permutation representations ρvert and ρdiag on vertices and diagonals respectively. Here’s what we have so
far:
rotations reflections
1 1 2 2 2
C1 C2 C3 C4 C5 Scalar product with self
Trivial χtriv 1 1 1 1 1 1
Orientation χori 1 1 1 −1 −1 1
Geometric χgeo 2 −2 0 0 0 1
Diagonal χdiag 2 2 0 2 0 2
Vertex χvert 4 0 0 2 0 3
Regular χreg 8 0 0 0 0
Evidently χtriv and χori are irreducible, and χgeo , χgeo = 1 so it must be the irreducible character of
dimension 2. Meanwhile, χdiag and χvert come from permutation representations, so they each must include
D E
at least one copy of the trivial character; in fact χdiag , χtriv = ⟨χvert , χtriv ⟩G = 1. This observation says
G
that we should look at the characters
χdiag − χtriv = [1, 1, −1, 1, −1], χvert − χtriv = [3, −1, −1, 1, −1].
The first of these is clearly irreducible; call it χ1 . The second one equals χ1 + χgeo , so this character does not
give us anything new.
To find the fifth and final irreducible character, one strategy is to use the regular character:
X
χreg = (dim χ)χ = χtriv + χori + 2χgeo + χ1 + χ2 .
irr. chars. χ
where χ2 is the irreducible character we are looking for. Solving this equation gives χ2 = [1, 1, −1, −1, 1].
Alternatively, we could have observed that χ2 = χ1 ⊗ χori . The final character table of D4 is as follows:
1 1 2 2 2
C1 C2 C3 C4 C5
χtriv 1 1 1 1 1
χori 1 1 1 −1 −1
χ1 1 1 −1 1 −1
χ2 1 1 −1 −1 1
χgeo 2 −2 0 0 0
In particular, each of the one-dimensional characters is its own dual, and in fact they form a Klein four-
group under tensor product, with duality as the inverse. In the next section, we take this observation and
run with it. ◀
A one-dimensional character of G is identical with the representation it comes from: a group homomor-
phism G → C× . Since tensor product is multiplicative on dimension, it follows that the tensor product of
two one-dimensional characters is also one-dimensional. In fact (χ ⊗ χ′ )(g) = χ(g)χ′ (g) (this is immediate
from the definition of tensor product) and χ ⊗ χ∗ = χtriv . So the set Ch(G) of one-dimensional characters,
i.e., Ch(G) = Hom(G, C× ), is an abelian group under tensor product (equivalently, pointwise multiplica-
tion), with identity χtriv .
194
Definition 8.8.1. The commutator of two elements a, b ∈ G is the element [a, b] = aba−1 b−1 . The (normal)
subgroup of G generated by all commutators is called the commutator subgroup, denoted [G, G]. The
quotient Gab = G/[G, G] is the abelianization of G.
The abelianization can be regarded as the group obtained by forcing all elements G to commute, in addition
to whatever relations already exist in G; in other words, it is the largest abelian quotient of G. It is routine
to check that [G, G] is indeed normal in G, and also that χ([a, b]) = 1 for all χ ∈ Ch(G) and a, b ∈ G. (In fact,
this condition characterizes the elements of the commutator subgroup, as will be shown soon.) Therefore
Ch(G) ∼= Ch(Gab ) and it suffices to understand the character groups of abelian groups.
Accordingly, let G be an abelian group of finite order n. The conjugacy classes of G are all singleton sets
(since ghg −1 = h for all g, h ∈ G), so there are n distinct irreducible representations of G. By (8.10) of
Theorem 8.6.5, so in fact every irreducible character is 1-dimensional (and every representation of G is a
direct sum of 1-dimensional representations). We have now reduced the problem to describing the group
homomorphisms G → C× .
The simplest case is that G = Z/nZ is cyclic. Write G multiplicatively, and let g be a generator. Then each
χ ∈ Ch(G) is determined by its value on g, which must be some nth root of unity. There are n possibilities
for χ(g), so all the irreducible characters of G arise in this way, and in fact form a group isomorphic to Z/nZ,
generated by any character that maps g to a primitive nth root of unity. So Hom(G, C× ) ∼ = G (although this
isomorphism is not canonical).
Now we consider the general case. Every abelian group G can be written as
r
G∼
Y
= Z/ni Z.
i=1
Let gi be a generator of the ith factor, and let ζi be a primitive (ni )th root of unity. Then each character χ is
determined by the numbers j1 , . . . , jr , where ji ∈ Z/ni Z and χ(gi ) = ζiji for all i. Thus Hom(G, C× ) ∼= G,
an isomorphism known as Pontryagin duality. More generally, for any finite group G (not necessarily
abelian), there is an isomorphism
Hom(G, C× ) ∼ = Gab . (8.11)
This is quite useful when computing the character table of a group: if you can figure out the commutator
subgroup and/or the abelianization, then you can immediately write down the one-dimensional characters.
Sometimes the size of the abelianization can be determined from the size of the group and the number of
conjugacy classes. (The commutator subgroup is normal, so it itself is a union of conjugacy classes.)
The description of characters of abelian groups implies that if G is abelian and g ̸= IdG , then χ(g) ̸= 1 for at
least one character χ. Therefore, for every group G, we have
195
If instead the group were known to have 6 conjugacy classes, then the equation has two solutions, namely
1,1,1,1,2,4 and 2,2,2,2,2,2, but the latter is impossible since every group has at least one 1-dimensional irrep,
namely the trivial representation. ◀
Example 8.8.3. Consider the case G = Sn . Certainly [Sn , Sn ] ⊆ An , and in fact equality holds. This is
trivial for n ≤ 2. If n ≤ 3, then the equation (a b)(b c)(a b)(b c) = (a b c) in Sn (multiplying left to right)
shows that [Sn , Sn ] contains every 3-cycle, and it is not hard to show that the 3-cycles generate the full
alternating group. Therefore (8.11) gives
Hom(Sn , C× ) ∼
= Sn /An ∼
= Z/2Z.
It follows that χtriv and χsign are the only one-dimensional characters of Sn . A more elementary way of
seeing this is that a one-dimensional character must map the conjugacy class of 2-cycles to either 1 or −1,
and the 2-cycles generate all of Sn , hence determine the character completely.
For instance, suppose we want to compute the character table of S5 (Problem 8.6), which has seven conju-
gacy classes. There are 21 lists of seven positive integers whose squares add up to |S5 | = 5! = 120, but only
four of them that contain exactly two 1’s:
1, 1, 2, 2, 2, 5, 9, 1, 1, 2, 2, 5, 6, 7, 1, 1, 2, 3, 4, 5, 8, 1, 1, 4, 4, 5, 5, 6.
By examining the defining representation and using the tensor product, you should be able to figure out
which one of these is the actual list of dimensions of irreps. ◀
Example 8.8.4. The dicyclic group G = Dic3 can be presented as
In particular it must have six irreps, four of dimension 1 and two of dimension 2. That’s a lot of 1’s, so
it is worth computing the commutator subgroup to get at the one-dimensional characters. It turns out
that [G, G] = {1, a2 , a4 }, and the quotient Gab is cyclic of order 4, generated by x. So the one-dimensional
characters are as follows:
1 1 2 2 3 3
C1 C2 C3 C4 C5 C6
χ1 1 1 1 1 1 1
χ2 1 1 1 1 −1 −1
χ3 1 −1 −1 1 i −i
χ4 1 −1 −1 1 −i i
The remaining two irreducible characters χ5 , χ6 evidently satisfy
196
Now let’s take some scalar products:
⟨χ1 , χ5 ⟩G = 2 + a + 2b + 2(−1 + c) = 0,
⟨χ3 , χ5 ⟩G = 2 − a − 2b + 2(−1 + c) = 0.
Adding these equations gives 4 + 4(−1 + c) = 0, or c = 0; subtracting them gives a = −2b. At this point
we know that the C3 and C4 columns of the character table are (1, 1, −1, −1, b, −b) and (1, 1, 1, 1, −1, 1)
respectively. By Proposition 8.6.7 they are orthogonal, i.e., 2 − 2b = 0, or b = 1. So the final character table
is as follows:
1 1 2 2 3 3
C1 C2 C3 C4 C5 C6
χ1 1 1 1 1 1 1
χ2 1 1 1 1 −1 −1
χ3 1 −1 −1 1 i −i
χ4 1 −1 −1 1 −i i
χ5 2 −2 1 −1 0 0
χ6 2 2 −1 1 0 0
◀
Let H ⊆ G be finite groups. Representations of G give rise to representations of H via an (easy) process
called restriction, and representations of H give rise to representations of G via a (somewhat more in-
volved) process called induction. These processes are sources of more characters to put in character tables,
and the two are related by an equation called Frobenius reciprocity.
Res(χρ ) = χρ |H = χ(Res(ρ)).
That is, restricting a representation does not change its character on the level of group elements. On the
other hand, the restriction of an irreducible representation is not always irreducible. Also, two elements
conjugate in G are not necessarily conjugate in a subgroup H, so restriction is not as trivial an operation as
it first might appear.
Example 8.9.1. Let Cλ denote the conjugacy class in Sn of permutations of cycle-shape λ. Recall that
the standard representation ρstd of G = S3 has character [2, 0, −1] on the conjugacy classes C111 , C21 , C3 .
Moreover, ⟨χstd , χstd ⟩G = 1 because ρstd is irreducible.
Now let H = A3 < S3 . This is an abelian group isomorphic to Z/3Z, so the two-dimensional representation
Res(ρstd ) cannot possibly be irreducible. In fact H = C111 ∪ C3 , so ⟨χstd , χstd ⟩G = (1 · 22 + 2 · (−1)2 )/3 = 2.
(We knew this already, since if a 2-dimensional representation is not irreducible then it must be the direct
sum of two one-dimensional irreps.) The group A3 is cyclic, so its character table is
IdG (1 2 3) (1 3 2)
χtriv 1 1 1
(8.12)
χ1 1 ω ω2
χ2 1 ω2 ω
197
where ω = e2πi/3 . (Note also that the conjugacy class C3 ⊆ S3 splits into two singleton conjugacy classes
in A3 .) Now it is evident that Res(χstd ) = [2, −1, −1] = χ1 + χ2 . ◀
V = CB ⊗ W = (b1 ⊗ W ) ⊕ · · · ⊕ (bn ⊗ W ).
To say how a group element g ∈ G acts on the summand bi ⊗ W , we need to write gbi in the form bj h, where
j ∈ [n] and h ∈ H. Note that for each bi , there is a unique pair bj , h that satisfies these conditions. We then
make g act by
g(bi ⊗ w) = bj ⊗ (h · w) = bj ⊗ ρ(h)(w), (8.13)
extended linearly to all of V . Heuristically, this formula is justified by the equation
In other words, g sends bi ⊗ W to bj ⊗ W , acting by h along the way. We have a map IndG
H (ρ) that sends each
g ∈ G to the linear transformation V → V just defined. Alternative notations for IndGH (ρ) include Ind(ρ) (if
G and H are clear from context) and ρ ↑G H .
Example 8.9.2. Let G = S3 and H = A3 = {Id, (1 2 3), (1 3 2)}, and let (ρ, W ) be a representation of H,
where W = C⟨e1 , . . . , en ⟩. Let B = {b1 = Id, b2 = (1 2)}, so that V = b1 ⊗ W ⊕ b2 ⊗ W . To define IndH
G (ρ),
we need to solve the equations gbi = bj h. That is, for each g ∈ G and each bi ∈ B, we need to determine the
unique pair bj , h that satisfy the equation.
i=1 i=2
g gbi = bj h gbi = bj h
Id Id = b1 Id (1 2) = b2 Id
(1 2 3) (1 2 3) = b1 (1 2 3) (1 3) = b2 (1 3 2)
(1 3 2) (1 3 2) = b1 (1 3 2) (2 3) = b2 (1 2 3)
(1 2) (1 2) = b2 Id Id = b1 Id
(1 3) (1 3) = b2 (1 2 3) (1 2 3) = b1 (1 2 3)
(2 3) (2 3) = b2 (1 3 2) (1 3 2) = b1 (1 3 2)
Therefore, the representation IndG H (ρ) sends the elements of S3 to the following block matrices. Each block
is of size n × n; the first block corresponds to b1 ⊗ W and the second block to b2 ⊗ W .
198
ρ(Id) 0 ρ(1 2 3) 0 ρ(1 3 2) 0
Id 7→ (1 2 3) 7→ (1 3 2) 7→
0 ρ(Id) 0 ρ(1 3 2) 0 ρ(1 2 3)
0 ρ(Id) 0 ρ(1 2 3) 0 ρ(1 3 2)
(1 2) 7→ (1 3) 7→ (2 3) 7→
ρ(Id) 0 ρ(1 2 3) 0 ρ(1 3 2) 0
For instance, if ρ is the 1-dimensional representation (= character) χ1 of (8.12), then the character of Ind(ρ)
is given on conjugacy classes in S3 by
χInd(ρ) (C111 ) = 2, χInd(ρ) (C21 ) = ω + ω 2 = −1, χInd(ρ) (C3 ) = 0,
Thus IndG
H (ρ) is as follows:
ρ(Id) 0 0 0 ρ(Id) 0 0 0 ρ(Id)
Id 7→ 0 ρ(Id) 0 (1 3) 7→ ρ(Id) 0 0 (2 3) 7→ 0 ρ(1 2) 0
0 0 ρ(Id) 0 0 ρ(1 2) ρ(Id) 0 0
ρ(1 2) 0 0 0 ρ(1 2) 0 0 0 ρ(1 2)
(1 2) 7→ 0 0 ρ(1 2) (1 2 3) 7→ 0 0 ρ(Id) (1 3 2) 7→ ρ(1 2) 0 0
0 ρ(1 2) 0 ρ(1 2) 0 0 0 ρ(Id) 0
◀
In fact, Ind(ρ) is a representation, and there is a general formula for its character. (That is a good thing,
because as you see computing the induced representation itself is a lot of work.)
Proposition 8.9.4. Let H be a subgroup of G and let (ρ, W ) be a representation of H with character χ. Then IndG
H (ρ)
is a representation of G, with character defined on g ∈ G by
1 X
IndG
H (χ)(g) = χ(k −1 gk).
|H|
k∈G: k−1 gk∈H
(We know that characters are class functions, so why not write χ(g) instead of χ(k −1 gk)? Because χ is a
function on H, so the former expression is not well-defined in general.)
199
Proof. First, we verify that Ind(ρ) is a representation. Let g, g ′ ∈ G and bi ⊗ w ∈ V . Then there is a unique
bk ∈ B and h ∈ H such that
gbi = bk h (8.15)
and in turn there is a unique bℓ ∈ B and h′ ∈ H such that
g ′ bk = bℓ h′ . (8.16)
Now that we know that Ind(ρ) is a representation of G on V , we calculate its character using (8.14):
n
X X
Ind(χ)(g) = tr(Bi,i ) = χ(b−1
i gbi )
i=1 b−1
i∈[n]: i gbi ∈H
X 1 X
= χ(h−1 b−1
i gbi h)
|H|
i∈[n]: b−1
i gbi ∈H
h∈H
1 X
= χ(k −1 gk) (8.17)
|H|
k∈G: k−1 gk∈H
as desired. Here k = bi h runs over all elements of G as the indices of summation i, h on the previous sum
run over [r] and H respectively. (Also, k −1 gk = h−1 b−1 −1
i gbi h ∈ H if and only if bi gbi ∈ H, simply because
H is a group.) Since Ind(χ) is independent of the choice of B, so is the isomorphism type of Ind(ρ).
Corollary 8.9.5. Let H ⊆ G and let ρ be the trivial representation of H. Then
#{k ∈ G : k −1 gk ∈ H}
IndG
H (χtriv )(g) = .
|H|
Proof. Normality implies that k −1 gk ∈ H if and only if g ∈ H, independently of k. If g ∈ H then the sum
in (8.17) has |G| terms, all equal to χ(g); otherwise, the sum is empty. (Alternative proof: normality implies
that left cosets and right cosets coincide, so the blocks in the block matrix Ind(ρ)(g) will all be on the main
diagonal (and equal to ρ(g)) when g ∈ H, and off the main diagonal otherwise.)
200
Example 8.9.7. Let n ≥ 2, so that An is a normal subgroup of Sn of index 2. By Corollary 8.9.5,
(
Sn 2 for g ∈ A3 ,
IndAn (χtriv )(g) =
0 for g ̸∈ A3 ,
#{k ∈ Sn | (k −1 gk)(n) = n}
IndS
Sn−1 (χtriv )(g) =
n
(n − 1)!
#{k ∈ Sn | k(n) is a fixed point of g}
=
(n − 1)!
= #{fixed points of g}
= χdef (g)
Example 8.9.9. Let G = S4 and let H be the non-normal subgroup {id, (1 2), (3 4), (1 2)(3 4)}. Let ρ be
the trivial representation of G and χ its character. We can calculate ψ = IndG
H (χ) using Corollary 8.9.5:
where, as usual, Cλ denotes the conjugacy class in S4 of permutations with cycle-shape λ. In the notation
of Example 8.7.3, the decomposition into irreducible characters is χ1 + χ2 + 2χ5 . ◀
Proof.
1 X
⟨Ind(χ), ψ⟩G = Ind(χ)(g) · ψ(g)
|G|
g∈G
1 X 1 X
= χ(k −1 gk) · ψ(g) (by Prop. 8.9.4)
|G| |H|
g∈G k∈G: k−1 gk∈H
1 XX X
= χ(h) · ψ(k −1 gk)
|G||H|
h∈H k∈G g∈G: k−1 gk=h
1 XX
= χ(h) · ψ(h) (i.e., g = khk −1 )
|G||H|
h∈H k∈G
1 X
= χ(h) · ψ(h) = ⟨χ, Res(ψ)⟩H .
|H|
h∈H
201
Example 8.9.11. Frobenius reciprocity sometimes suffices to calculate the isomorphism type of an induced
representation. Let G = S3 and H = A3 , and let χstd , χ1 and χ2 be as in Example 8.9.1. We would like to
compute Ind(χ1 ). By Frobenius reciprocity
But χstd is irreducible and dim χstd = dim Ind(χ1 ) = 2. Therefore, it must be the case that Ind(χ1 ) = χstd ,
and the corresponding representations are isomorphic. The same is true if we replace χ1 with χ2 . ◀
We have worked out the irreducible characters of S3 , S4 and S5 ad hoc (the last as an exercise). In fact, we
can do this for all n, exploiting a vast connection to the combinatorics of partitions and tableaux.
Recall (Defn. 1.2.4) that a partition of n is a sequence λ = (λ1 , . . . , λℓ ) of weakly decreasing positive integers
whose sum is n. We will sometimes drop the parentheses and commas. We write λ ⊢ n or |λ| = n to indicate
that λ is a partition of n. The number ℓ = ℓ(λ) is the length of λ. The set of all partitions of n is Par(n), and
the number of partitions of n is p(n) = |Par(n)|. For example,
We will write Par for the set of all partitions. (As a set this is the same as Young’s lattice.)
For each λ ⊢ n, let Cλ be the conjugacy class in Sn consisting of all permutations with cycle shape λ. Since
the conjugacy classes are naturally indexed by Par(n), it makes sense to look for a set of representations
indexed by partitions.
Definition 8.10.1. Let µ = (µ1 , . . . , µℓ ) ⊢ n.
• The Ferrers diagram of shape µ is the top- and left-justified array of boxes with µi boxes in the ith
row.
• A (Young) tableau6 of shape µ is a Ferrers diagram with the numbers 1, 2, . . . , n placed in the boxes,
one number to a box.
• Two tableaux T, T ′ of shape µ are row-equivalent, written T ∼ T ′ , if the numbers in each row of T
are the same as the numbers in the corresponding row of T ′ .
• A (Young) tabloid of shape µ is an equivalence class of tableaux under row-equivalence. A tabloid
can be represented as a tableau without vertical lines.
• We write sh(T ) = µ to indicate that a tableau or tabloid T is of shape µ.
1 3 6 1 3 6
2 7 2 7
4 5 4 5
increase downward and rightward. In these notes, I will call such a tableau a “standard tableau”. For the moment, I am not placing
any restrictions on which numbers can go in which boxes: thus there are n! Young tableaux of shape µ for any µ ⊢ n.
202
A Young tabloid can be regarded as an ordered set partition (T1 , . . . , Tm ) of [n] in which |Ti | = µi . The
order of the blocks Ti matters, but not the order of entries within each block. Thus the number of tabloids
of shape µ is
n n!
= .
µ µ1 ! · · · µm !
The symmetric group Sn acts on tabloids by permuting the numbers. This action gives rise to a permutation
representation (ρµ , Vµ ) of Sn , the µ-tabloid representation of Sn . Here Vµ is the vector space of all formal
C-linear combinations of tabloids of shape µ. The character of ρµ will be denoted τµ .
Example 8.10.2. For n = 3, the characters of the tabloid representations ρµ are as follows.
Conjugacy classes
C111 C21 C3
τ3 1 1 1 (8.18)
Characters τ21 3 1 0
τ111 6 0 0
|Cµ | 1 3 2
◀
ρ(n) ∼
= ρtriv .
• The tabloids of shape µ = (1, 1, . . . , 1) are just the permutations of [n]. Therefore
ρ(1,1,...,1) ∼
= ρreg .
• A tabloid of shape µ = (n − 1, 1) is determined by its singleton part. So the representation ρµ is
isomorphic to the action on the singleton by permutation, i.e.,
ρ(n−1,1) ∼
= ρdef .
In fact, all tabloid representations can be obtained as inductions of an appropriate trivial representation.
Some notation first. For a partition λ = (λ1 , . . . , λℓ ) ⊢ n, define
so that Sλ ∼ = Sλ1 × · · · × Sλℓ . This is called a Young subgroup; its cardinality is λ!. For example, if
λ = (4, 4, 2, 1) ⊢ 11, then Sλ is the subgroup of S11 mapping each of the sets {1, 2, 3, 4}, {5, 6, 7, 8}, {9, 10},
{11} to itself. Note that Sλ is not a normal subgroup of Sn unless λ = (n) or λ = (1n ), because replacing
L1 , . . . , Lℓ with any partition of [n] into blocks of the same sizes gives a conjugate subgroup.
Proposition 8.10.3. Let λ = (λ1 , . . . , λℓ ) ⊢ n. Then IndS ∼
Sλ (ρtriv ) = ρλ .
n
Proof. We will show that the characters of these representations are equal, i.e., that IndS
Sλ (χtriv ) = τλ .
n
Assign labels 1, . . . , n to the cells of the Ferrers diagram of λ reading from left to right and top to bottom, so
that the cells in the ith row are labeled by Li . For every w ∈ Sn , let Tλ,w be the tableau of shape λ in which
cell k is filled with the number w(k).
203
Let g be a permutation of cycle type µ = (µ1 , . . . , µk ), say
and let Mj = [µ[j−1] + 1, µ[j] ] be the support of the jth cycle of g. Then, by Corollary 8.9.5 (replacing w with
w−1 for brevity),
1
IndS
Sλ (χtriv )(g) =
n
#{w ∈ Sn : wgw−1 ∈ Sλ }
λ!
1 n o
= # w ∈ Sn : w(1) · · · w(µ1 ) · · · w(n − µk + 1) · · · w(n) ∈ Sλ
λ!
1
= # {w ∈ Sn : ∀j ∃i : w(Mj ) ⊆ Li }
λ!
1
= # {tableaux Tλ,w with all elements of Mj in the same row}
λ!
= # {tabloids of shape λ with all elements of Mj in the same row}
= # {tabloids of shape λ fixed by g}
= τλ (g)
= τλ (Cµ ).
For n = 3, the table in (8.18) is a triangular matrix. In particular, the characters τµ are linearly independent,
hence a basis, in the vector space Cℓ(S3 ). We will prove that this is the case for all n. We will need two
orders on the set Par(n).
Definition 8.10.4. Let λ, µ ∈ Par(n).
2. Dominance order: λ ◁ µ (“λ is dominated by µ”) if λ ̸= µ and λ[k] ≤ µ[k] for all k.
(5) > (4, 1) > (3, 2) > (3, 1, 1) > (2, 2, 1) > (2, 1, 1, 1) > (1, 1, 1, 1, 1).
(“Lex-greater partitions are short and wide; lex-smaller ones are tall and skinny.”)
Dominance is a partial order on Par(n). It first fails to be a total order for n = 6 (neither of 33 and 411
dominates the other). Lex order is a linear extension of dominance order: if λ ◁ µ then λ < µ.
Since the tabloid representations ρµ are permutation representations, we can calculate τµ by counting fixed
points. That is, for any permutation w ∈ Cλ , we have
1. τµ (Cµ ) ̸= 0.
2. τµ (Cλ ) ̸= 0 only if λ ⊴ µ (thus, only if λ ≤ µ in lex order).
Proof. First, let w ∈ Cµ . Take T to be any tabloid whose blocks are the cycles of w; then wT = T . For
example, if w = (1 3 6)(2 7)(4 5) ∈ S7 , then T can be either of the following two tabloids:
204
1 3 6 1 3 6
2 7 4 5
4 5 2 7
Q
It follows from (8.21) that τµ (Cµ ) ̸= 0. In fact, τµ (Cµ ) = j rj !, where rj is the number of occurrences of j
in µ.
For the second assertion, observe that w ∈ Sn fixes a tabloid T of shape µ if and only if every cycle of w
is contained in a row of T . This is possible only if, for every k, the largest k rows of T are collectively big
enough to hold the k largest cycles of w. This is precisely the condition λ ⊴ µ.
Dominance is not enough for τµ (Cλ ) to be nonzero: for instance, take λ = (2, 2) and µ = (3, 1).
Proof. Make the characters into a p(n) × p(n) matrix X = [τµ (Cλ )]µ,λ⊢n with rows and columns ordered by
lex order on Par(n). By Proposition 8.10.5, X is a triangular matrix with nonzero entries on the diagonal, so
it is nonsingular.
We can transform the characters τµ into a list of irreducible characters χµ of Sn by applying the Gram-
Schmidt process with respect to the inner product ⟨·, ·⟩Sn . We will start with µ = (n) and work our way
up in lex order. In fact, the τµ are not only a vector space basis for Cℓ(Sn ), but an integral basis for the
Z-module ClZ (Sn ) of virtual characters (although we cannot prove that statement at this point). Thus no
division will be required in the Gram-Schmidt process. Each tabloid character τµ will decompose as
X
τµ = χµ + Kλ,µ χλ (8.22)
λ<µ
for certain integers Kλ,µ called Kostka numbers. We will have more to say about them in the next chapter.
For the time being, we will only be able to observe (8.22) for particular examples, including S3 and S4 , but
we will eventually be able to prove it in general (Corollary 9.11.3). The irrep with character χµ is called a
Specht module. There is an independent construction of Specht modules, which I have not written up yet.
Example 8.10.7. We will use tabloid representations to derive the character tables of S3 and S4 .
Recall the table of characters (8.18) of the tabloid representations for G = S3 . Here is how the Gram-
Schmidt process goes.
Second, ⟨τ21 , χ3 ⟩G = 1. Thus τ21 − χ3 = [2, 0, −1] is orthogonal to χ3 , and in fact it is irreducible, so we
label it as χ21 . (This is χstd .)
205
which is 1-dimensional, hence irreducible (in fact it is χsign ). To summarize,
X
K
z }| {
z }| { "C111 C21 C3
τ3 1 1 1 1 0 0 1 1 1 # χ3
τ21 = 3 1 0 = 1 1 0 2 0 −1 χ2 (8.23)
τ111 6 0 0 1 2 1 1 −1 1 χ111
The matrix K has 1’s on the main diagonal, which witnesses the earlier claim that the tabloid characters
form an integral basis for CℓZ (Sn ). ◀
At this point you should feel a bit dissatisfied, since I have not told you what the values of the irreducible
characters actually are in general, just that you can obtain them from the tabloid characters plus Gram-
Schmidt. That is a harder problem; the answer is given by the Murnaghan–Nakayama Rule (see §9.14), which
expresses the values of the irreducible characters as signed counts of certain tableaux.
I have also not told you the multiplicities of the irreps in the tabloid representations, i.e., the numbers in
the matrix K. For S3 and S4 these matrices are unitriangular (1’s on the main diagonal and 0’s above); in
fact this property holds for all Sn . Thus the tabloid characters are not just a vector basis for Cℓ(Sn ), but,
more strongly, a basis for the free abelian group generated by irreducible characters. We will prove this
eventually (Corollary 9.11.3), by which point we will have a combinatorial description of the entries of K.
8.11 Exercises
In all exercises, unless otherwise specified, G is a finite group and (ρ, V ) and (ρ′ , V ′ ) are finite-dimensional
representations of G over C.
Problem 8.1. Let χ be an irreducible character of G and let ψ be a one-dimensional character. Prove that
ω := χ ⊗ ψ is an irreducible character.
206
Problem 8.2. Let n ≥ 2. Prove that the standard representation ρstd of Sn (see Example 8.3.4) is irreducible.
(Hint: Compute ⟨χdef , χdef ⟩ and ⟨χdef , χtriv ⟩. The latter boils down to finding the expected number of fixed
points in a permutation selected uniformly at random; this is an old classic that uses what is essentially
linearity of expectation.)
1 X
⟨χdef , χdef ⟩Sn = #{i ∈ [n] : σ(i) = i}2
n!
σ∈Sn
1 X
= #{(i, j) ∈ [n]2 : σ(i) = i, σ(j) = j}
n!
σ∈Sn
1 X
= #{σ ∈ Sn : σ(i) = i, σ(j) = j}
n!
(i,j)∈[n]2
1 X X
= #{σ ∈ Sn : σ(i) = i} + #{σ ∈ Sn : σ(i) = i, σ(j) = j}
n!
i∈[n] i,j∈[n]
i̸=j
1
= (n · (n − 1)! + (n(n − 1)) · (n − 2)!)
n!
= 2.
It follows that ρdef must be the direct sum of two non-isomorphic irreps. One of them must be ρtriv because
we know that the all-1’s vector is fixed by permutations. Alternatively, we can find the multiplicity of ρtriv
explicitly:
1 X
⟨χtriv , χdef ⟩Sn = #{i ∈ [n] : σ(i) = i}
n!
σ∈Sn
(
1 X X 1 if σ(i) = i
=
n! 0 if σ(i) ̸= i
σ∈Sn i∈[n]
(
1 X X 1 if σ(i) = i
=
n! 0 if σ(i) ̸= i
i∈[n] σ∈S n
1 X
= #{σ ∈ Sn : σ(i) = i}
n!
i∈[n]
1 X
= (n − 1)!
n!
i∈[n]
n · (n − 1)
= = 1.
n!
(Alternatively, observe that ⟨χdef , χtriv ⟩Sn is evidently positive, so it must be 1 by what we know about the
decomposition of ρdef into irreps.) So ρstd = ρdef − ρtriv is the other irreducible summand. Alternatively,
⟨χstd , χstd ⟩Sn = ⟨χdef , χdef ⟩Sn − ⟨χdef , χtriv ⟩Sn − ⟨χtriv , χdef ⟩Sn + ⟨χtriv , χtriv ⟩Sn = 2 − 1 − 1 + 1 = 1.
207
Problem 8.3. Let G be a group of order 63. Prove that G cannot have exactly 5 conjugacy classes. (You are
encouraged to use a computer for part of this problem.)
Solution: If such a group exists, then the list of dimensions of its irreps must be a positive integer solution
P5
to the equation i=1 d2i = 63. The following Sage code will produce all solutions:
Solution: (a) Here is a table of representatives of the five conjugacy classes of S4 and their actions on X,
with fixed points indicated in boldface. As usual, the value of the character χρ is the number of fixed points.
For later use, I also include the character of ρdef and the size of each conjugacy class.
Permutation Id (1 2) (1 2 3) (1 2)(3 4) (1 2 3 4)
Conj. class size 1 6 8 3 6
12|34 12|34 12|34 14|23 12|34 14|23
13|24 13|24 14|23 12|34 13|24 13|24
14|23 14|23 13|24 13|24 14|23 12|34
χρ (σ) 3 1 0 3 1
χdef (σ) 4 2 1 0 0
(b) As for all permutation representations, the one-dimensional subspace W = C⟨12|34 + 13|24 + 14|23⟩ is
G-invariant and carries the trivial action. The character of G on the invariant complement W ⊥ is
and this character is irreducible because its scalar product with itself is 1 (calculation omitted).
208
Problem 8.5. Prove that the irreps of a direct product G × G′ are exactly the direct products (see Exam-
ple 8.1.11) of the irreps of G with irreps of G′ . (This was mentioned in §8.8 for abelian groups but in fact is
true in general.)
Solution: Suppose (ρ, V ) and (ρ′ , V ′ ) are representations, with characters χ and χ′ respectively.. Then there
is a representation ρ × ρ′ of G × G′ on V ⊗ V ′ given by
(ρ × ρ′ )(g, g ′ ) = ρ(g) ⊗ ρ′ (g ′ )
or equivalently
(ρ × ρ′ )(g, g ′ )(v ⊗ v ′ ) = ρ(g)(v) ⊗ ρ′ (g ′ )(v ′ )
with character χρ×ρ′ (g, g ′ ) = χρ (g)χρ′ (g), so that
1 X
χρ×ρ′ , χρ×ρ′ = χρ×ρ′ (g, g ′ )χρ×ρ′ (g, g ′ )
G×G′ |G × G′ |
(g,g ′ )∈G×G′
1 X X
= ′
χρ (g)χρ′ (g ′ )χρ (g)χρ′ (g ′ )
|G| |G |
g∈G g ′ ∈G′
1 X 1 X
= χρ (g)χρ (g) ′ χ ′ (g ′ )χρ′ (g ′ )
|G| |G | ′ ′ ρ
g∈G g ∈G
= χρ , χρ G
χρ′ , χρ′ G′
.
In particular, if ρ, ρ′ are both irreps then so is ρ × ρ′ . The conjugacy classes in G × G′ are just products of
conjugacy classes in G and G′ , so we have constructed the right number of irreps. Alternatively, we can see
that we have a complete list by observing that if ρ1 , . . . , ρn and ρ′1 , . . . , ρ′m are complete rosters of irreps for
G and G′ respectively, then
!
X X X X
dim(ρi × ρ′j )2 = (dim ρi )2 (dim ρ′j )2 = (dim ρi )2 (dim ρ′j )2 = |G| |H| = |G × H|.
i,j i,j i j
Problem 8.6. Work out the character table of S5 without using any of the material in Section 8.10. (Hint: To
construct another irreducible character, start by considering the action of S5 on the edges of the complete
graph K5 induced by the usual permutation action on the vertices.)
Solution: There are 7 conjugacy classes, corresponding to the 7 possible cycle shapes of a permutation,
as given by the partitions of 5. Call them C1 , C2 , C22 , C3 , C32 , C4 , C5 . Their sizes are 1, 10, 15, 20, 20, 30, 24
respectively. We will write each character as a bracketed list of its values on these conjugacy classes in order.
We start with the irreducible characters
Also, χ4 = χ3 ⊗ χsign = [4, −2, 0, 1, 1, 0, −1] is irreducible (either by Problem 8.1, or by checking directly
that ⟨χ4 , χ4 ⟩ = 1).
So far we have four irreducible characters of dimensions 1, 1, 4, 4. Therefore, the squares of the dimensions
of the remaining three irreducibles must add up to 120 − 1 − 1 − 16 − 16 = 86. The only possibilities are 5,
5, 6.
209
The standard permutation action of S5 on [5] induces an action on the edges of the complete graph K5 ,
giving a 10-dimensional representation with character
χedge = [10, 4, 2, 1, 1, 0, 0].
We have
χedge , χedge = 3, χedge , χtriv = 1, χedge , χ3 = 1
(calculation omitted), which says that
χ5 = χedge − χtriv − χ3 = [5, 1, 1, −1, 1, −1, 0]
is irreducible, and so is χ6 = χ5 ⊗ χsign .
The last irreducible character χ7 must have dimension 6, so we can solve for it using the regular character:
6
!
1 X
χ7 = χreg − (dim χi )χi .
6 i=1
Problem 8.7. Work out the character table of the quaternion group Q; this is the group of order 8 whose
elements are {±1, ±i, ±j, ±k} with relations i2 = j 2 = k 2 = −1, ij = k, jk = i, ki = j.
Solution: The center is {1, −1}. Note that j −1 = −j and so jij −1 = −jij = −jk = −i. So i and −i are
conjugate, and similar calculations show that the conjugacy classes are {1}, {−1}, {±i}, {±j}, {±k}. We
need five numbers whose squares add up to 8; the only possibility is 1,1,1,1,2. The commutator subgroup is
{1, −1} (this is not hard to check directly), so the abelianization is a Klein-four group, letting us write down
the four one-dimensional characters. At this point I will actually write down the full character table, which
is as follows:
Character 1 −1 ±i ±j ±k
χ1 1 1 1 1 1
χ2 1 1 1 −1 −1
χ3 1 1 −1 1 −1
χ4 1 1 −1 −1 1
χ5 2 −2 0 0 0
The two-dimensional irreducible character can be found by solving the equation
4
X
χreg = [8, 0, 0, 0, 0] = χj + 2χ5 .
j=1
210
Problem 8.8. Work out the irreducible characters of S5 using tabloid characters. Feel free to use a computer
algebra system to automate the tedious parts. Compare your result to the character table of S5 calculated
ad hoc in Problem 8.6. Make as many observations or conjectures as you can about how the partition λ is
related to the values of the character χλ , and about the Kostka numbers.
Solution: As usual, let τµ denote the character of the tabloid representation ρµ , and let Cλ denote the
conjugacy class of cycle-shape λ. The tabloid characters are given by the following table:
Conjugacy classes
1 10 15 20 20 30 24
C11111 C2111 C221 C311 C32 C41 C5
τ5 1 1 1 1 1 1 1
τ41 5 3 1 2 0 1 0
τ32 10 4 2 1 1 0 0
Tabloid characters τ311 20 6 0 2 0 0 0
τ221 15 3 2 0 0 0 0
τ2111 60 6 0 0 0 0 0
τ11111 120 0 0 0 0 0 0
Doing Gram-Schmidt provides the Kostka numbers and the irreducible characters:
X
K
z }| {
z }| { C11111 C2111 C221 C311 C32 C41 C5
τ5 1 0 0 0 0 0 0 1 1 1 1 1 1 1
χ5
τ41 1 1 0 0 0 0 0 4 2 0 1 −1 0 −1 χ41
τ32 1 1 1 0 0 0 0
5 1 1 −1 1 −1 0
χ32
τ311 = 1 2 1 1 0 0 0 6 0 −2 0 0 0 1
χ311 .
τ221 1 2 2 1 1 0 0
5 −1 1 −1 −1 1 0
χ221
τ2111 1 3 3 3 2 1 0 4 −2 0 1 1 0 −1 χ2111
τ11111 1 4 5 6 5 4 1 1 −1 1 1 −1 −1 1 χ11111
• The transition matrix is unitriangular, so the tabloid characters are a Z-basis for the space of integer-
valued class functions.
• dim χλ is the number of standard Young tableaux of shape λ.
• Conjugating (transposing) the partition corresponds to tensoring with χsign .
Problem 8.9. Recall that the alternating group An consists of the n!/2 even permutations in Sn , that is, those
with an even number of even-length cycles.
(a) Show that the conjugacy classes in A4 are not simply the conjugacy classes in S4 . (Hint: Consider the
possibilities for the dimensions of the irreducible characters of A4 .)
(b) Determine the conjugacy classes in A4
(c) Use this information to determine [A4 , A4 ] without computing any more commutators.
211
(d) Now compute the character table of A4 .
Solution: (a) We have A4 = C1111 ∪C31 ∪C22 as a union of conjugacy classes in S4 , and any two permutations
conjugate in PA4 are certainly conjugate in S4 , so there are at least three conjugacy classes. On the other hand,
|A4 | = 12 = ρ (dim ρ)2 , where ρ runs over all irreps in A4 , and the equation 12 = 1 + x2 + y 2 has no integer
solutions, so the number of conjugacy classes cannot be three.
(b) The identity is in a conjugacy class by itself. Also (multiplying left to right),
must be contained in conjugacy classes, hence must be conjugacy classes. Note that 12 can be written as the
sum of four squares: 9+1+1+1, which says that the irreps of A4 must have dimensions 3, 1, 1, 1.
(c) Let K = [A4 , A4 ] be the commutator subgroup. The one-dimensional characters of the abelianization
A4 /K are the three one-dimensional characters of A4 . By Pontryagin duality we have |A4 /K| = 3, so
|K| = 12/3 = 4, and the only possibility is K = C1111 ∪ C22 , which is isomorphic to a Klein four-group.
(d) The group of one-dimensional characters has size 3. Letting ω = e2πi/3 , it must be
′ ′′
Character 1 C22 C31 C31
χ1 1 1 1 1
χ2 1 1 ω ω2
χ3 1 1 ω2 ω
212
Chapter 9
Symmetric Functions
The symmetric polynomials that are homogeneous of degree d form a finitely generated, free R-module
Λd (R). For example, if n = 3, then up to scalar multiplication, the only symmetric polynomial of degree 1
in x, y, z is x + y + z. In degree 2, here are two:
x2 + y 2 + z 2 , xy + xz + yz.
Every other symmetric polynomial that is homogeneous of degree 2 is a R-linear combination of these
two, because the coefficients of x2 and xy determine the coefficients of all other monomials. Similarly, the
polynomials
x3 + y 3 + z 3 , x2 y + xy 2 + x2 z + xz 2 + y 2 z + yz 2 , xyz
are a basis for the space of degree 3 symmetric polynomials in R[x, y, z].
Each member of this basis is a sum of the monomials in a single orbit under the action of S3 . Accordingly,
we can index them by the partition whose parts are the exponents of one of its monomials. That is,
m3 (x, y, z) = x3 + y 3 + z 3 ,
m21 (x, y, z) = x2 y + xy 2 + x2 z + xz 2 + y 2 z + yz 2 ,
m111 (x, y, z) = xyz.
But this sum is empty is if ℓ > n. So if we want to construct a basis for the symmetric polynomials indexed
by partitions, n variables is not enough — we need a countably infinite set of variables {x1 , x2 , . . . }, which
means that we need to work not with polynomials, but with formal power series.
213
9.2 Formal power series
Let R be an integral domain (typically Z or a field), and let xQ= {x1 , x2 , . . . } be a countably infinite Pset of
∞
commuting indeterminates. A monomial is a product xα = i=1 xα i
i
, where α i ∈ N for all i and i∈I αi
is finite (equivalently, all but finitely many of the αi are zero). The sequence α = (α1 , α2 , . . . ) is called the
exponent vector of the monomial; listing the nonzero entries of α in decreasing order gives a partition λ(α).
A formal power series is an expression X
cα xα
α
with cα ∈ R for all α. Equivalently, a formal power series can be regarded as a function from monomials
to R, mapping xα to cα . (Also equivalently, the R-module of all formal power series is thus the direct
product (not the direct sum) of infinitely many copies of R, one for each exponent vector.) We often use the
notation
[xα ]F = coefficient of monomial xα in the power series F .
The set R[[x]] of all formal power series is an abelian group under addition, and in fact an R-module,
namely the direct product of countably infinitely many copies of R, one for each exponent vector. 1 In fact,
R[[x]] is a ring as well, with multiplication given by
!
X X X X
cα xα dβ x β = cα dβ xγ .
α∈NI β∈NI γ∈NI (α,β): α+β=γ
because the inner sum on the right-hand side has only finitely many terms for each γ, and is thus a well-
defined element of R.
We are generally not concerned with whether (or where) a formal power series converges in the sense of
calculus, since we rarely need to plug in real values for the indeterminates xi (and when we do, analytic
convergence is not usually an issue). All that matters is that every operation must produce a well-defined
power series, in the sense that each coefficient is given by a finite computation in the base ring R. For
example, multiplication of power series satisfies this criterion, as explained above.2
Familiar functions from analysis (like exp and log) can be regarded as formal power series, namely their
Taylor series. However, we will typically study them using combinatorial rather than analytic methods.
For instance, from this point of view, we would justify equating the function 1/(1 − x) as equal to the
power series 1 + x + x2 + · · · not by calculating derivatives of 1/(1 − x), but rather by observing that the
identity (1 − x)(1 + x + x2 + · · · ) = 1 holds in Z[[x]]. (That said, combinatorics also gets a lot of mileage
out of working with derivative operators — but treating them formally, as linear transformations that map
monomials to other monomials, rather than analytically.) Very often, analytical identities among power
series can be proved using combinatorial methods; see Problem 9.6 for an example.
We can now define symmetric functions properly, as elements of the ring of formal power series C[[x]] =
C[[x1 , x2 , . . . ]].
1 By contrast, the polynomial ring R[x] is the direct sum of countably infinitely many copies of R.
2 We would have a problem with multiplication if we allowed two-way-infinite series. For example, the square of n∈Z xn is not
P
well-defined.
214
Definition 9.3.1. Let λ ⊢ n. The monomial symmetric function mλ is the power series
X
mλ = xα
α: λ(α)=λ
where, as before, α runs over all infinite lists of nonnegative integers in which all but finitely many entries
are zero.
For example,
∞
X
m3 = x31 + x32 + x33 + · · · = x3i ,
i=1
X
m21 = x21 x2 + x1 x22 + x21 x3 + x1 x32 + x22 x3 + x2 x23 + · · · = x2i xj ,
i̸=j
X
m111 = x1 x2 x3 + x1 x2 x4 + x1 x3 x4 + x2 x3 x4 + x1 x2 x5 + · · · = xi xj xk .
1≤i<j<k
where
Each Λd is a finitely generated free R-module, with basis {mλ : λ ⊢ d}, and their direct sum Λ is a graded R-
algebra. If we let S∞ be the group whose members
S∞ are the permutations of {x1 , x2 , . . . } with only finitely
many non-fixed points (equivalently, S∞ = n=1 Sn ), then Λ is the ring of formal power series that have
bounded degree and that are invariant under the action of S∞ .
The monomial symmetric functions are the most obvious basis for Λ from an algebraic point of view, in
the sense that each mλ is the orbit under S∞ of any single monomial in it. On the other hand, there are
many other bases that arise more frequently in combinatorics. Understanding symmetric functions requires
familiarity with these various bases and how they interact.
One piece of terminology: we say that a basis B of Λ is an integral basis if the symmetric functions with
integer coefficients are precisely the integer linear combinations of elements of B. Evidently, {mλ } is an
integral basis. This condition is stronger than being a vector space basis for Λ; for example, integral bases
are not preserved by scaling.
Definition 9.4.1. The kth elementary symmetric function ek is the sum of all squarefree monomials of
degree k. That is,
e0 = 1,
X Y X
ek = xs = x i1 x i2 · · · x ik for k > 0,
S⊆N>0 s∈S 0<i1 <i2 <···<ik
|S|=k
215
Equivalently, ek = m1k , where 1k means the partition with k 1′ s. For λ = (λ1 , . . . , λℓ ) ∈ Par, we define
eλ = eλ1 · · · eλℓ .
In general, we say that a basis for Λ is multiplicative if it is defined on partitions in this way. (Note that
{mλ } is not multiplicative, which is why Problem 9.1 is nontrivial.)
Apparently {e3 , e21 , e111 } is an R-basis for Λ3 , because the transition matrix is unitriangular and therefore
invertible over every R. This works for n = 4 as well, where
e1111 24 12 6 4 1 m1111
e211 12 5 2 1 0 m211
e22 = 6 2 1 0 0 m22 .
e31 4 1 0 0 0 m31
e4 1 0 0 0 0 m4
This matrix is again unitriangular, and notably is symmetric across the northwest/southeast diagonal —
that is, the coefficient of eλ in mµ equals the coefficient of eµ in mλ .
## Input
n = 3
e = SymmetricFunctions(QQ).elementary()
m = SymmetricFunctions(QQ).monomial()
for lam in Partitions(n):
m(e[lam])
## Output
m[1, 1, 1]
3*m[1, 1, 1] + m[2, 1]
6*m[1, 1, 1] + 3*m[2, 1] + m[3]
Let ⊵ denote the dominance partial order on partitions (see Definition 8.10.4). Also, for a partition λ, let λ̃
be its conjugate, given by transposing the Ferrers diagram (see the discussion after Example 1.2.4).
Theorem 9.4.2. Let λ, µ ⊢ n, with ℓ = ℓ(λ) and k = ℓ(µ). Let bλ,µ be the coefficient3 of eλ when expanded in the
monomial basis, that is, X
eλ = bλ,µ mµ .
µ
Then bλ,λ̃ = 1, and bλ,µ = 0 unless λ̃ ⊵ µ. In particular, {eλ : λ ⊢ n} is an integral basis for Λn .
3 Stanley [Sta99, §7.4] writes Mλ,µ .
216
Proof. Say that a λ-factorization of a monomial is a factorization into monomials of degrees λ1 , . . . , λℓ . Let
xµ = xµ1 1 · · · xµk k . Then
Represent such a λ-factorization of xµ by a tableau T of shape λ in which the ith row contains the variables
in xαi , in increasing order. We will say such a tableau has content µ; i.e., its entries consist of µ1 1’s, µ2 2’s,
etc.
For example, suppose that µ = (3, 2, 2, 1, 1) and λ = (4, 2, 2, 1). One λ-factorization of xµ and its associated
tableau are
x31 x22 x23 x14 x15 = (x1 x2 x3 x5 )(x1 x3 )(x2 x4 )(x1 ), T = 1 2 3 5 .
1 3
2 4
1
Thus the entries of T correspond to variables, and its rows correspond to factors. Observe that all the 1’s in
T must be in the first column; all the 2’s must be in the first or second column; etc. Thus, for every j, there
must be collectively enough boxes in the first j columns of T to hold all the entries of T corresponding to
the variables x1 , . . . , xj . That is,
which is precisely the condition λ̃ ⊵ µ. If this fails, then no λ-factorization of xµ can exist and bλ,µ = 0.
If λ̃ = µ, then every inequality in (9.1) is in fact an equality, which says that every entry in the jth column
is in fact j. That is, there is exactly one λ-factorization of xµ , and bλ,µ = 1.
Therefore, if we order partitions of n by any linear extension of dominance (such as lex order), then the
matrix [bλ,µ ] will be upper unitriangular, hence invertible over any integral domain R. (This is the same
argument as in Corollary 8.10.6.) It follows that the R-module spanned by the eλ ’s is the same as that
spanned by the mµ ’s for any R, so {eλ } is an integral basis.
Corollary 9.4.3 (“Fundamental Theorem of Symmetric Functions”). The elementary symmetric functions e1 , e2 , . . .
are algebraically independent. Therefore, Λ = R[e1 , e2 , . . . ] as rings.
Proof. Given any nontrivial polynomial relation among the ei ’s, extracting the homogeneous pieces would
give a nontrivial linear relation among the eλ ’s, which does not exist.
Corollary 9.4.4. The transition matrix between the bases {eλ } and {mµ } is symmetric; that is, bλ,µ = bµ,λ .
Proof. We know from the proof of Theorem 9.4.2 that bλ,µ = [mµ ]eλ is the number of λ-factorizations of xµ
into squarefree monomials xα1 , . . . , xαℓ . Moreover, these factorizations are in bijection with row-increasing
tableaux T of shape λ and content µ: Specifically, xαj is the product of the variables xi such that i appears
in the jth row of T . On the other hand, such tableaux are also in bijection with µ-factorizations of xλ into
squarefree monomials xβ1 , . . . , xβk , where xβj is the product of the variables xi such that j appears in the
ith row of T . The conclusion follows.
217
For instance, if µ = (3, 2, 2, 1, 1) and λ = (4, 2, 2, 1), then the tableau
1 2 3 5
1 3
2 4
1
Definition 9.5.1. The kth complete homogeneous symmetric function hk is the sum of all monomials of
degree k, extended multiplicatively to partitions:
h0 = 1,
X X
hk = xi1 xi2 · · · xik = mλ for k > 0,
0<i1 ≤i2 ≤···≤ik λ⊢k
The coefficient matrices above are all Z-invertible, witnessing the fact that the hλ ’s are also a free R-module
basis for Λ and that Λ = R[h1 , h2 , . . . , ] as a ring. We could figure out the transition matrices between the hλ
and eλ , but instead will take a different approach that exploits the close relation between the two families.
Consider the generating functions
X X
E(t) = tk ek , H(t) = tk hk .
k≥0 k≥0
We regard E(t) and H(t) as formal power series in t whose coefficients are themselves formal power series
in {xi }. Observe that
X Y X Y 1
E(t) = tk ek = (1 + txi ), H(t) = tk hk = . (9.2)
1 − txi
k≥0 i≥1 k≥0 i≥1
In the formula for H(t), each factor in the infinite product is a geometric series 1+txi +t2 x2i +· · · , so [tk ]H(t)
is the sum of all monomials of degree k. Now (9.2) implies that
∞ X
X n
H(t)E(−t) = (−1)k ek hn−k tk = 1
n=0 k=0
218
and extracting the coefficients of positive powers of t gives the Jacobi-Trudi relations: for every n ≥ 0,
n
X
(−1)k ek hn−k = 0 ∀n > 0. (9.3)
k=0
That is,
h1 − e1 = 0, h2 − e1 h1 + e2 = 0, h3 − e1 h2 + e2 h1 − e3 = 0, ...
(where we have plugged in h0 = e0 = 1). The Jacobi-Trudi relations can be used iteratively to solve for the
ek in terms of the hk :
e1 = h1 ,
e2 = e1 h1 − h2 = h21 − h2 ,
(9.4)
e3 = e2 h1 − e1 h2 + h3 = h1 (h21 − h2 ) − h2 h1 + h3 = h31 − 2h1 h2 + h3 ,
e4 = e3 h1 − e2 h2 + e1 h3 − h4 = h41 − 3h21 h2 + h22 + 2h1 h3 − h4 ,
etc. Since the Jacobi-Trudi relations are symmetric in the letters h and e, so are the equations (9.4). Therefore,
the elementary and homogenous functions generate the same ring.
Here is another way to see that the h’s are an integral basis, which again exploits the symmetry of the
Jacobi-Trudi relations. Define a ring endomorphism ω : Λ → Λ by
ω(ei ) = hi (9.5)
for all i, extended multiplicatively (so that ω(eλ ) = hλ ) and linearly to all symmetric functions. This map,
sometimes known as the Hall transformation4 but more usually just referred to as ω, is well-defined since
the elementary symmetric functions are algebraically independent. Now Corollary 9.5.2 follows from the
following result:
Proof. Applying ω to the Jacobi-Trudi relations (9.3), we see that for every n ≥ 1,
n
X n
X
0 = (−1)n−k ω(ek )ω(hn−k ) = (−1)n−k hk ω(hn−k )
k=0 k=0
Xn
= (−1)k hn−k ω(hk ) (by replacing k with n − k)
k=0
n
X
= (−1)n (−1)n−k hn−k ω(hk )
k=0
and comparing this last expression with the original Jacobi-Trudi relations gives ω(hk ) = ek (e.g., because
solving for ω(hk ) in terms of the hk ’s gives exactly (9.4), with the ek ’s replaced by ω(hk )’s).
219
9.6 Power-sum symmetric functions
Definition 9.6.1. The kth power-sum symmetric function pk is the sum of the kth powers of all variables,
extended multiplicatively to partitions:
∞
X
pk = mk = xki ,
i=1
pλ = pλ1 · · · pλℓ for λ = (λ1 , . . . , λℓ ) ∈ Par.
Note that the transition matrices are invertible over Q, but not over Z: for example, m11 = (p11 − p2 )/2.
Thus the power-sums are not an integral basis of symmetric functions (although, as we will shortly prove,
they are a vector space basis for ΛQ ).
We have seen this transition matrix before: its columns are characters of tabloid representations! (See (8.23).)
This is the first explicit connection we can observe between representations of Sn and symmetric functions,
and it is the tip of an iceberg. It is actually not hard to prove.
Theorem 9.6.2. For λ ⊢ n, we have X
pλ = τµ (Cλ )mµ
µ⊢n
where τµ (Cλ ) means the character of the tabloid representation of shape µ on the conjugacy class Cλ of cycle-shape λ,
as in §8.10.
Let λ = (λ1 , . . . , λℓ ) and µ = (µ1 , . . . , µk ). We adopt the notation (8.19). As in Theorem 9.4.2, let
Proof. Q
xµ = i xµi i . We calculate the coefficient
Here we will represent each such choice by a tabloid T in which the factor xλcii contributes labels Li to the
ci th row, so that T has shape µ. Thus the rows of T correspond to variables, while the entries correspond to
positions in the factorization (in contrast to the construction of Theorem 9.4.2).
For example, let λ = (2, 1, 1, 1) and µ = (3, 2). Then [xµ ]pλ = [x31 x22 ]pλ = 4. The four λ-factorizations of xµ
are shown below with their corresponding tabloids.
These are precisely the tabloids in which each interval Li is contained in a single row, and these are precisely
those fixed by the permutation given in cycle notation as
220
whose cycle-shape is λ. (Compare equation (8.20). In the example above, w is the transposition (1 2).) In
particular, the number of such tabloids is by definition τµ (Cλ ).
Corollary 9.6.3. {pλ } is a basis for the symmetric functions (although not an integral basis).
Proof. By Proposition 8.10.5, the transition matrix [τµ (Cλ )] from the monomial symmetric functions to the
power-sums is triangular, hence invertible (although not unitriangular).
The definition of Schur symmetric functions is very different from the m’s, e’s, h’s and p’s. It is not even
clear at first that they are symmetric. But in fact the Schur functions turn out to be essential in the study of
symmetric functions and in several ways are the “best” basis for Λ.
Definition 9.7.1. A column-strict tableau T of shape λ, or λ-CST for short, is a labeling of the boxes of the
Ferrers diagram of λ with integers (not necessarily distinct) that is
The partition λ is called the shape of T , and the set of all column-strict tableaux of shape λ is denoted
CST(λ). The content of a CST is the sequence α = (α1 , α2 , . . . ), where αi is the number of boxes labelled i,
and the weight of T is the monomial xT = xα = xα 1 α2
1 x2 · · · (the same information as the content, but in
monomial form). For example:
1 1 3 1 1 1 1 2 3
2 3 4 8 1 4
The terminology is not entirely standardized; column-strict tableaux are often called “semistandard tableaux”
(as in, e.g. [Sta99]).
Definition 9.7.2. The Schur function corresponding to a partition λ is
X
sλ = xT .
T ∈CST(λ)
At the other extreme, suppose that λ = (1, 1, . . . , 1) is the partition with n singleton parts, so that the
corresponding Ferrers diagram has a single column. To construct a CST of this shape, we need n distinct
221
labels, which can be arbitrary. Therefore
◀
Example 9.7.4. Let λ = (2, 1). We will express sλ as a sum of the monomial symmetric functions m3 , m21 , m111 .
First, no tableau of shape λ can have three equal entries, so the coefficient of m3 is 0.
Second, for weight xa xb xc with a < b < c, there are two possibilities, shown below.
a b a c
c b
Third, for every a ̸= b ∈ N>0 , there is one tableau of shape λ and weight x2a xb : the one on the left if a < b,
or the one on the right if a > b.
a b b b
b a
It should be evident at this point that the Schur functions are quasisymmetric, i.e., that for every monomial
xai11 · · · xaikk (where i1 < · · · < ik ), its coefficient in sλ depends only on the ordered sequence (a1 , . . . , ak ). To
see this, observe that if j1 < · · · < jk , then replacing is with js for all s gives a bijection from λ-CSTs with
weight xai11 · · · xaikk to λ-CSTs with weight xaj11 · · · xajkk .
In fact, the Schur functions are symmetric. Here is an elementary proof. It is enough to show that sλ is
invariant under transposing xi and xi+1 for every i ∈ N>0 , since those transposition generate S∞ . Let
T ∈ CST(λ) and consider all the entries equal to i or i + 1, ignoring columns that contain both i and i + 1.
For ease in depicting the the set of such entries in a single row looks like
Say that there are a instances of i and b instances of i + 1. Then we can replace this part of the tableau with b
instances of i and a instances of i + 1. Doing this for every row gives a shape-preserving bijection between
tableaux of weight · · · xpi xqi+1 · · · and those of weight · · · xqi xpi+1 · · · , as desired.
An important generalization of a Schur function involves a generalization of the underlying Ferrers dia-
gram of a tableau.
Definition 9.7.5. Let λ, µ be partitions with µ ⊆ λ, i.e., λi ≥ µi for all i. There is then an associated skew
partition or skew shape λ/µ, defined via its skew Ferrers diagram, in which the ith row has boxes in
columns µi + 1, . . . , λi . A skew tableau of shape λ/µ is a filling of the skew Ferrers diagram with numbers.
222
Some skew shapes are shown below; note that disconnected skew shapes are possible.
The notion of a column-strict tableau carries over without change to skew shapes. Here is a CST of shape
λ/µ, where λ = (8, 6, 6, 5, 4, 2) and µ = (5, 3, 3, 3, 2):
2 2 3
1 1 3
2 3 4
4 4
2 6
1 1
The definition of Schur functions (Definition 9.7.1) can also be adapted to skew shapes.
Definition 9.7.6. Let CST(λ/µ) denote the set of all column-strict skew tableaux of shape λ/µ, and as before
Q α (T )
weight each tableau T ∈ CST(λ/µ) by the monomial xT = i xi i , where αi (T ) is the number of i’s in T .
The skew Schur function is then X
sλ/µ = xT .
T ∈CST(λ/µ)
The elementary proof of symmetry of Schur functions carries over literally to skew Schur functions.
We are next going to establish a formula for the Schur function sλ as a determinant of a matrix whose entries
are hn ’s or en ’s (which also proves their symmetry). This takes more work, but the proof, due to the ideas
of Lindström, Gessel, and Viennot, is beautiful, and the formula has many other useful consequences (such
as what the involution ω does to Schur functions). This exposition follows closely that of [Sag01, §4.5].
Theorem 9.8.1. For any λ = (λ1 , . . . , λℓ ) we have
sλ = det hλi −i+j i,j=1,...,ℓ (9.8)
and
sλ̃ = det eλi −i+j i,j=1,...,ℓ . (9.9)
In particular, the Schur functions are symmetric.
For example,
h3 h4 h5 h3 h4 h5
s311 = h0 h1 h2 = 1 h1 h2 = h311 + h5 − h41 − h32 .
h−1 h0 h1 0 1 h1
223
(6, ∞)
6
5
4
4
3
3 3
2
1
1
0
0 1 2 3 4 5 6 7 8
Figure 9.1: A lattice path P from (1, 0) to (6, ∞) with weight xP = x1 x23 x4 x6 .
Proof. We prove (9.8) in detail, and then discuss how the proof can be modified to prove (9.9).
We will consider lattice paths P that start at some point on the x-axis in Z2 and move north or east one unit
at a time. For every path that we consider, the number of eastward steps must be finite, but the number of
northward steps is infinite. Thus the “ending point” is (x, ∞) for some x ∈ N. Label each eastward Q step e
of P by the number L(e) that is its y-coordinate plus one. The weight of P is the monomial xP = e xL(e) .
An example is shown in Figure 9.1.
The monomial xP determines the path P up to horizontal shifting, and xP can be any monomial. Thus we
have a bijection, and it follows that for any a ∈ N,
X X
hn = xP = xP . (9.10)
paths P from paths P with fixed starting
(a, 0) to (a + n, ∞) point with n east steps
Step 2: Express the generating function for families of lattice paths in terms of the hk ’s.
For a partition λ of length ℓ, a λ-path family P = (π, P1 , . . . , Pℓ ) consists of the following data:
• A permutation π ∈ Sℓ ;
• Two sets of points U = {u1 , . . . , uℓ } and V = {v1 , . . . , vℓ }, defined by
Figure 9.2 shows a λ-path family with λ = (3, 3, 2, 1) and π = 3124. (In general the paths in a family are
allowed to share edges, although that is not the case in this example.)
224
v4 v3 v2 v1
4
P4 P3 P2 P1
3
1
u4 u3 u2 u1
0
0 1 2 3 4 5 6 7 8
Note that for each i ∈ [ℓ], the number of east steps in the path Pi from u(π(i)) to vi is
Now the first miracle occurs: the signed generating function for path families is the determinant of a matrix
whose entries are complete homogeneous symmetric functions! One key observation is that any collection
of paths P1 , . . . , Pℓ in which Pi contains λi − i + π(i) east steps gives rise to a λ-path family (π, P1 , . . . , Pℓ ).
In other words, if we know what π is, then Pi can be any path with the appropriate number of east steps.
Qℓ
For a path family P = (π, P1 , . . . , Pℓ ), let xP = i=1 xPi and (−1)P be the sign of π. Then:
X X X
(−1)P xP = ε(π) xP1 · · · xPℓ
P=(π,P1 ,...,Pℓ ) π∈Sℓ λ-path families
P=(π,P1 ,...,Pℓ )
X ℓ
Y X
xP i
= ε(π)
(by the key observation above)
π∈Sℓ i=1 paths Pi with
λi −i+π(i) east steps
X ℓ
Y
= ε(π) hλi −i+π(i) (by (9.10))
π∈Sℓ i=1
Call a path family good if no two of its paths meet in a common vertex, and bad otherwise. Note that if P
is good, then π must be the identity permutation, and in particular (−1)P = 1.
225
Define a sign-reversing, weight-preserving involution P 7→ P♯ on bad λ-path families as follows.
1. Of all the lattice points contained in two or more paths in P, choose the point α with the lex-greatest
pair of coordinates.
2. Of all the half-paths from α to some vi , choose the two with the largest i. Interchange them. Call the
resulting path family P♯ .
v4 v3 v2 v1 v4 v3 v2 v1
5 5
4 4
3 P 3 P]
2 2
α α
1 1
u4 u3 u2 u1 u4 u3 u2 u1
0 0
0 1 2 3 4 5 6 0 1 2 3 4 5 6
Then
For each good path family, label the east steps of each path by height as before. The labels weakly increase
as we move north along each path. Moreover, for every j the jth east step of the path Pi occurs one unit
east of that of Pi+1 , so it must also occur strictly north of it (otherwise, the paths would cross). Therefore,
we can construct a column-strict tableau of shape λ by reading off the labels of each path, and this gives
a bijection between good λ-path families and column-strict tableaux of shape λ. An example is shown in
Figure 9.4.
Consequently, (9.12) implies that |hλi −i+j |i,j=1,...,ℓ = sλ , which is (9.8). Is that amazing or what?
The proof of (9.9) is similar. The key difference is that instead of labeling each east step with its height,
we number all the steps (north and east) consecutively, ignoring the first i − 1 steps of Pi (those below the
226
v4 v3 v2 v1
1 1 3
4 4
2 3 4
3 4
3 3 3
5
1 1
u4 u3 u2 u1
Figure 9.4: The bijection between good path families and column-strict tableaux.
line y = x + ℓ − 1, which must all be northward anyway). The weight of a path is still the the product
of the variables corresponding to its east steps. This provides a bijection between lattice paths with k east
steps and squarefree monomials of degree k, giving an analogue of (9.10), with hn replaced by en . Bad path
families cancel out by the same involution as before, and each good path family now gives rise to a tableau
of shape λ in which rows strictly increase but columns weakly increase (see Figure 9.5). Transposing gives
a column-strict tableau of shape λ̃, and (9.9) follows.
Corollary 9.8.2. For every partition λ, the involution ω interchanges sλ and sλ̃ .
Proof. We know that ω interchanges hλ and eλ , so it interchanges the RHS’s, hence the LHS’s, of (9.8)
and (9.9).
The next step is to prove that the Schur functions are a basis for the symmetric functions. Now that we
know they are symmetric, they can be expressed in the monomial basis as
X
sλ = Kλ,µ mµ . (9.13)
µ⊢n
Thus Kλ,µ is the number of column-strict tableaux T with shape λ and content µ. These are called the
Kostka numbers.
Theorem 9.8.3. The Schur functions {sλ : λ ⊢ n} are a Z-basis for ΛZ .
Proof. Here comes one of those triangularity arguments. Consider the matrix of Kostka numbers [Kλ,µ ]λ,µ⊢n .
First, if λ = µ, then there is exactly one possibility for T : fill the ith row full of i’s. Therefore
∀λ ⊢ n : Kλ,λ = 1. (9.14)
Second, observe that if T is a CST of shape λ and content µ (so in particular Kλ,µ > 0), then
227
v4 v3 v2 v1
1 1 2 5
3 5
1 3 5
2 4 1 3
1 3 5
2
2 4
1
3
1 2
u4 u3 u2 u1
Figure 9.5: The dual bijection between good path families and row-strict tableaux.
The lattice-path proof of Theorem 9.8.1 generalizes to skew shapes (although I haven’t yet figured out
exactly how) to give Jacobi-Trudi determinant formulas for skew Schur functions:
sλ/µ = det hλi −µi −i+j i,j=1,...,ℓ , sλ̃/µ̃ = det eλi −µi −i+j i,j=1,...,ℓ . (9.15)
The next step in studying the ring of symmetric functions Λ will be to define an inner product structure
on it. These will come from considering the Cauchy kernel and the dual Cauchy kernel, which are formal
power series in two sets of variables x = {x1 , x2 , . . . }, y = {y1 , y2 , . . . }, defined as the following infinite
products: Y Y
Ω= (1 − xi yj )−1 , Ω∗ = (1 + xi yj ).
i,j≥1 i,j≥1
∗
The power series Ω and Ω are well-defined because the coefficient of any monomial xα yβ is the number
of ways of factoring it into monomials of the form xi yj , which is clearly finite (in particular it is zero if
|α| ̸= |β|). Moreover, they are evidently bisymmetric5 , i.e., symmetric with respect to each of the variable
5 Technically, Ω lives not in the ring Λ(x, y) of bisymmetric power series, but rather its completion, since it contains terms of
arbitrarily high degree. If you don’t know what “completion” means then don’t worry about it. The key point is that Ω is still
determined by the coefficients of the bisymmetric series uλ (x)vµ (y) for any bases {uλ }, {vµ } of Λ — it is just no longer true that all
but finitely many of these coefficients are zero.
228
sets x = {x1 , x2 , . . . } and y = {y1 , y2 , . . . }. Thus we can write Ω and Ω∗ as power series in some basis for
Λ(x) and ask which elements of Λ(y) show up as coefficients.
For example, if λ = (3, 3, 2, 1, 1, 1) then zλ = (13 3!)(21 1!)(32 2!) = 216 and ελ = −1.
Proposition 9.9.1. Let λ ⊢ n and let Cλ be the corresponding conjugacy class in Sn . Then |Cλ | = n!/zλ , and ελ is
the sign of each permutation in Cλ .
as desired. (Regard λ as the partition whose parts are α1 , . . . , αℓ , sorted in weakly decreasing order.)
For the second equality in (9.17), recall the standard power series expansions
X qn X qn X qn
log(1 + q) = (−1)n+1 , log(1 − q) = − , exp(q) = . (9.19)
n n n!
n≥1 n≥1 n≥0
6 In [Sta99], Stanley uses m where I use r , presumably as a mnemonic for “multiplicity.” I have changed the notation in order to
i i
avoid conflict with the notation for monomial symmetric functions.
229
Q P
These are formal power series that obey the rules you would expect; for instance, log( i qi ) = i (log qi )
and exp log(q) = q. (The proof of the second of these is left to the reader as Problem 9.6.) In particular,
Y X
log Ω = log (1 − xi yj )−1 = − log(1 − xi yj )
i,j≥1 i,j≥1
X X xni yjn
= (by (9.19))
n
i,j≥1 n≥1
X1 X X X pn (x)pn (y)
= xni yjn =
n n
n≥1 i≥1 j≥1 n≥1
and now exponentiating both sides and applying the power series expansion for exp, we get
k
X pn (x)pn (y) X 1 X pn (x)pn (y)
Ω = exp =
n k! n
n≥1 k≥0 n≥1
X 1 r1 r2
X k! p1 (x)p1 (y) p2 (x)p2 (y)
= ···
k! r1 ! r2 ! · · · 1 2
k≥0 λ: ℓ(λ)=k
i=1
The proofs of the identities for the dual Cauchy kernel are analogous, and are left to the reader as Prob-
lem 9.7.
As a first benefit, we can express the homogeneous and elementary symmetric functions in the power-sum
basis.
Corollary 9.9.3. For all n, we have:
P
1. hn = λ⊢n pλ /zλ ;
P
2. en = λ⊢n ελ pλ /zλ ;
3. ω(pλ ) = ελ pλ (where ω is the involution of 9.5).
Set y1 = t, and yk = 0 for all k > 1. This kills all terms on the left side for which λ has more than one part,
leaving only those where λ = (n), while on the right side pλ (y) specializes to t|λ| , so we get
X X pλ (x)t|λ|
hn (x)tn =
n
zλ
λ
230
and extracting the coefficient of tn gives the desired expression for hn .
(3) Let ω act on symmetric functions in x while fixing those in y. Using (9.17) and (9.18), we obtain
! !
X pλ (x) pλ (y) X X X pλ (x)pλ (y)
= hλ (x)mλ (y) = ω eλ (x)mλ (y) = ω ελ
zλ zλ
λ λ λ λ
X ελ ω(pλ (x)) pλ (y)
=
zλ
λ
and equating the red coefficients of pλ (y)/zλ yields the desired result.
Definition 9.9.4. The Hall inner product on symmetric functions is defined by declaring {hλ } and {mλ } to
be dual bases. That is, we define
⟨hλ , mµ ⟩Λ = δλµ
and extend by linearity to all of Λ.
Thus the Cauchy kernel can be regarded as a generating function for pairs (hλ , mµ ), weighted by their inner
product. In fact it can be used more generally to compute Hall inner products:
Proposition 9.9.5. The Hall inner product has the following properties:
P
1. If {uλ } and {vµ } are graded bases for Λ indexed by partitions, such that Ω = λ uλ (x)vλ (y), then they are
dual bases with respect to the Hall inner product; i.e., ⟨uλ , vµ ⟩ = δλµ .
√
2. In particular, {pλ } and {pλ /zλ } are dual bases, and {pλ / zλ } is self-dual, i.e., orthonormal.
3. ⟨·, ·⟩ is a genuine inner product (in the sense of being a nondegenerate bilinear form).
4. ⟨·, ·⟩ is positive-definite: ⟨f, f ⟩ ≥ 0 for all f ∈ Λ, with equality if and only if f = 0.
5. The involution ω is an isometry with respect to the Hall inner product, i.e.,
Proof. Assertion (1) is a matter of linear algebra, and is left to the reader (Problem 9.3). Assertion (2) follows
from (1) together with (9.18), and (3) from the fact that ΛR admits an orthonormal basis, namely {pλ /zλ }.
Expanding an arbitrary symmetric function f in the power-sum basis quickly yields (4). Finally, the quickest
proof of (5) uses the power-sum basis: by Corollary 9.9.3(3), we have
√
The orthonormal basis {pλ / zλ } is not particularly nice from a combinatorial point of view, because it
involves irrational coefficients. It turns out that there is a better orthonormal basis: the Schur functions!
Our next goal will be to prove that
Y 1 X
Ω = = sλ (x)sλ (y) (9.20)
1 − xi yj
i,j≥1 λ
231
9.10 The Robinson-Schensted-Knuth correspondence
Recall from Example 1.2.4 that a standard [Young] tableau of shape λ is a filling of the Ferrers diagram of
λ with the numbers 1, 2, . . . , n that is increasing left-to-right and top-to-bottom. We write SYT(λ) for the set
of all standard tableaux of shape λ, and set f λ = |SYT(λ)| (this symbol f λ is traditional).
For example, if λ = (3, 3), then f λ = 5; the members of SYT(λ) are as follows.
1 3 5 1 3 4 1 2 5 1 2 4 1 2 3
2 4 6 2 5 6 3 4 6 3 5 6 4 5 6
• If T = ∅, then T ← x = x .
• If x ≥ u for all entries u in the top row of T , then append x to the end of the top row.
• Otherwise, find the leftmost entry u such that x < u. Replace u with x, and then insert u into the
subtableau consisting of the second and succeeding rows. In this case we say that x bumps u.
• Repeat until the bumping stops.
Step 1: Row-insert w1 = 5 into P . We do this in the obvious way. Since it is the first cell added, we add a
cell containing 1 to Q.
P = 5 Q= 1 (9.21a)
Step 2: Row-insert w2 = 7 into P . Since 5 < 7, we can do this by appending the new cell to the top row, and
adding a cell labeled 2 to Q to record where we have put the new cell in P .
P = 5 7 Q= 1 2 (9.21b)
Step 3: Row-insert w3 = 2 into P . This is a bit trickier. We cannot just append a 2 to the first row of P ,
because the result would not be a standard tableau. The 2 has to go in the top left cell, but that already
7 Different versions of the algorithm are referred to by various subsets of these three names; I am not drawing that distinction.
232
contains a 5. Therefore, the 2 “bumps” the 5 out of the first row into a new second row. Again, we record
the location of the new cell by adding a cell labeled 3 to Q.
2 7 1 2
P = Q= (9.21c)
5 3
Step 4: Row-insert w4 = 1 into P . This time, the new 1 bumps the 2 out of the first row. The 2 has to go into
the second row, but again we cannot simply append it to the right. Instead, the 2 bumps the 5 out of the
second row into the (new) third row.
1 7 1 2
P = Q= (9.21d)
2 3
5 4
Step 5: Row-insert w5 = 4 into P . The 4 bumps the 7 out of the first row. The 7, however, can comfortably
fit at the end of the second row, without any more bumping.
P = 1 4 Q= 1 2 (9.21e)
2 7 3 5
5 4
Step 6: Row-insert w6 = 8 into P . The 8 just goes at the end of the first row.
1 4 8 1 2 6
P = Q= (9.21f)
2 7 3 5
5 4
P = 1 3 8 Q= 1 2 6 (9.21g)
2 4 3 5
5 7 4 7
P = 1 3 6 Q= 1 2 6 (9.21h)
2 4 8 3 5 8
5 7 4 7
◀
233
A crucial feature of the RSK correspondence is that it can be reversed. That is, given a pair (P, Q), we can
recover the permutation that gave rise to it.
Example 9.10.3. Suppose that we are given the pair of tableaux in (9.21h). What was the previous step? To
get the previous Q, we just delete the 8. As for P , the last cell added must be the one containing 8. This is in
the second row, so somebody must have bumped 8 out of the first row. That somebody must be the largest
number less than 8, namely 6. So 6 must have been the number inserted at this stage, and the previous pair
of tableaux must have been those in (9.21g). ◀
Example 9.10.4. Let P be the standard tableau (with 18 boxes) shown in (a) below. Suppose that we know
that the cell labeled 16 was the last one added (because the corresponding cell in Q contains an 18). Then
the “bumping path” must be as indicated in the center figure (b). (That is, the 16 was bumped by the 15,
which was bumped by the 13, and so on.) Each number in the bumping path is the rightmost one in its row
that is less than the next lowest number in the path. The previous tableau in the RSK algorithm can now
be found by “unbumping”: push every number in the bumping path up and toss out the top one, to obtain
the tableau on the right (c).
1 2 5 8 10 18 1 2 5 8 10 18 1 2 5 8 12 18
(a) (b) (c)
3 4 11 12 19 3 4 11 12 19 3 4 11 13 19
6 7 13 6 7 13 6 7 15
9 15 17 9 15 17 9 16 17
14 16 14 16 14
Iterating this procedure allows us to recover w from the pair (P, Q). ◀
1 2 3 1 2 1 3 1
3 2 2
3
So
f (4) = 1, f (3,1) = 3, f (2,2) = 2, f (2,1,1) = 3, f (1,1,1,1) = 1.
and the sum of the squares of these numbers is 24. ◀
234
We have seen these numbers before — they are the dimensions of the irreps of S3 and S4 , as calculated in
Examples 8.7.2 and 8.7.3. Hold that thought!
The proof is in [Sta99, §7.13]; I hope to understand and write it up some day. It is certainly not obvious
from the standard RSK algorithm, where it looks like P and Q play inherently different roles. In fact, they
are more symmetric than they look. There are alternative descriptions of RSK from which the symmetry is
more apparent, also described in [Sta99, §7.13] and in [Ful97, §4.2]. I describe the former (without proof) in
§9.18.
The RSK correspondence can be extended to more general tableaux. This turns out to be the key to expand-
ing the Cauchy kernel in terms of Schur functions.
Definition 9.10.10. A generalized permutation of length n is a 2 × n matrix
q q1 q2 · · · qn
w= = (9.22)
p p1 p2 · · · pn
where q = (q1 , . . . , qn ), p = (p1 , . . . , pn ) ∈ Nn>0 , and the (q1 , p1 ), . . . , (qn , pn ) are in lex order. (That is, q1 ≤
· · · ≤ qn , and if qi = qi+1 then pi ≤ pi+1 .) The weight of w is the monomial xP yQ = xp1 · · · xpn yq1 · · · yqn .
The set of all generalized permutations will be denoted GP, and the set of all generalized permutations of
length n will be denoted GP(n).
q
If qi = i for all i and the pi ’s are pairwise distinct elements of [n], then w = p is just an ordinary permuta-
tion in Sn , written in two-line notation.
The generalized RSK algorithm (gRSK) is defined in exactly the same way as original RSK, except that
the inputs are now allowed to be generalized permutations rather than ordinary permutations. At the ith
stage, we row-insert pi in the insertion tableau P and place qi in the recording tableau Q in the new cell
added.
Example 9.10.11. Consider the generalized permutation
q 1 1 2 4 4 4 5 5 5
w= = ∈ GP(9).
p 2 4 1 1 3 3 2 2 4
The result of the gRSK algorithm is as follows. The unnamed tableau on the right records the step in which
each box was added.
P = 1 1 2 2 4 Q= 1 1 4 4 5 1 2 5 6 9
2 3 3 2 4 5 3 4 8
4 5 7
◀
The tableaux P, Q arising from gRSK will always have the same shape as each other, and will be weakly
increasing eastward and strictly increasing southward — that is, they will be column-strict tableaux, precisely
the things for which the Schur functions are generating functions. Column-strictness of P follows from the
235
definition of insertion. As for Q, it is enough to show that no label k appears more than once in the same
column. Indeed, all instances of k in q occur consecutively (say as qi , . . . , qj ), and the corresponding entries
of p are weakly increasing, so none of them will bump any other (in fact their bumping paths will not cross),
which means that each k appears to the east of all previous k’s.
This observation also suffices to show that the generalized permutation w can be recovered from the pair
(P, Q): the rightmost instance of the largest entry in Q must have been the last box added. Hence the
corresponding box of P can be “unbumped” to recover the previous P and thus the last column of w.
Iterating this process allows us to recover w. Therefore, generalized RSK gives a bijection
RSK
[
GP(n) −−−→ {(P, Q) : P, Q ∈ CST(λ)} (9.23)
λ⊢n
q
maps to a pair of tableaux P, Q with weight monomials xP and yQ .
in which a generalized permutation p
On the other hand, a generalized permutation w = pq ∈ GP(n) is determined by the number aij of occur-
rences of every column pqii . Therefore, the generating function for generalized permutations by weights is
for every λ. So the two blue expressions are equal. Applying ω to both sides (using Prop. 9.9.5(5)) we get
* +
X
⟨sλ , eµ ⟩ = ω(sλ ), ω(eµ ) = ⟨sλ̃ , hµ ⟩ = Kλ̃,µ = sλ , Kλ̃µ sλ (9.25)
λ
236
9.11 The Frobenius characteristic
As in Section 8.6, denote by Cℓ(Sn ) the vector space of C-valued class functions on the symmetric group Sn ;
also, let Cℓ(S0 ) = C. Define a graded vector space
M
Cℓ(S) = Cℓ(Sn )
n≥0
We now want to make Cℓ(S) into a graded ring. To start, we declare that the elements of Cℓ(S0 ) behave
like scalars. For n1 , n2 ∈ N>0 and fi ∈ Cℓ(Sni ), we would like to define a product f1 f2 ∈ Cℓ(Sn ), where
n = n1 + n2 . First, define a function f1 × f2 : Sn1 × Sn2 → C by
this is a class function because the conjugacy classes in G × H are just the Cartesian products of conjugacy
classes in G with those in H (this is a general fact about products of groups). The next step is to lift to
Sn . Identify Sn1 × Sn2 with the Young subgroup Sn1 ,n2 ⊆ Sn fixing each of the sets {1, 2, . . . , n1 } and
{n1 + 1, n1 + 2, . . . , n1 + n2 }. (See (8.19).) We now define the product f1 f2 ∈ Cℓ(Sn ) by the formula for
induced characters (Proposition 8.9.4):
1 X
f1 f2 = IndSn
Sn (f1 × f2 ) = (f1 × f2 )(g −1 wg).
1 ,n2 n1 ! n2 !
g∈Sn :
g −1 wg∈Sn1 ,n2
There is no guarantee that f1 f2 is a character of Sn (unless f1 and f2 are characters), but at least this oper-
ation is a well-defined map on class functions, and it makes Cℓ(S) into a commutative graded C-algebra.
(It is pretty clearly bilinear and commutative; it is nontrivial but not hard to check that it is associative.)
For a partition λ ⊢ n, let 1λ ∈ Cℓ(Sn ) be the indicator function on the conjugacy class Cλ ⊆ Sn , and let
For a permutation w ∈ Sn , let λ(w) denote the cycle-shape of w (so λ(w) is a partition). Define a function
ψ : Sn → Λn by
ψ(w) = pλ(w) . (9.26)
Note that ψ is a class function (albeit with values in Λ rather than in C).
Definition 9.11.1. The Frobenius characteristic is the map
ch : CℓC (S) → ΛC
defined on f ∈ Cℓ(Sn ) by
1 X X pλ
ch(f ) = ⟨f, ψ⟩Sn = f (w) pλ(w) = f (Cλ )
n! zλ
w∈Sn λ⊢n
1. ch(1λ ) = pλ /zλ .
237
2. ch is an isometry, i.e., it preserves inner products:
3. ch is a ring isomorphism.
4. ch(IndS Sλ χtriv ) = ch(τλ ) = hλ .
n
5. ch(IndS Sλ χsign ) = eλ .
n
6. Let χ be any character of Sn and let χsign be the sign character on Sn . Then ch(χ ⊗ χsign ) = ω(ch(χ)), where
ω is the involution of 9.5.
7. ch restricts to an isomorphism CℓV (S) → ΛZ , where CℓV (S) is the Z-module generated by irreducible
characters (i.e., the space of virtual characters).
8. The irreducible characters of Sn are {ch−1 (sλ ) : λ ⊢ n}.
Proof. (1): Immediate from the definition. It follows that ch is (at least) a graded C-vector space isomor-
phism, since {1λ : λ ⊢ n} and {pλ /zλ : λ ⊢ n} are C-bases for Cℓ(Sn ) and Λn respectively.
1 X 1
⟨1λ , 1µ ⟩Sn = 1λ (w)1µ (w) = |Cλ |δλµ = δλµ /zλ = ⟨pλ /zλ , pµ /zµ ⟩Λ = ch(1λ ), ch(1µ ) Λ
n! n!
w∈Sn
where the penultimate equality is (9.17) (from expanding the Cauchy kernel in the power-sum bases).
(3): Let n = j + k and let f ∈ Cℓ(S[j] ) and g ∈ Cℓ(S[j+1,n] ) (so that elements of these two groups commute,
and the cycle-type of a product is just the multiset union of the cycle-types). Then:
D E
ch(f g) = IndS n
Sj ×Sk (f × g), ψ (where ψ is defined as in (9.26))
S
D E n
= f × g, ResS Sj ×Sk ψ
n
(by Frobenius reciprocity)
Sj ×Sk
1 X
= f × g(w, x) · pλ(wx)
j! k!
(w,x)∈Sj ×Sk
!
1 X 1 X
= f (w) pλ(w) g(x) pλ(x) (because the power-sum basis is multiplicative)
j! k!
w∈Sj x∈Sk
= ch(f ) ch(g).
(4), (5): Denote by χntriv and χnsign the trivial and sign characters on Sn . We calculate in parallel:
1 X 1 X
= pλ(w) = ελ(w) pλ(w) (by def’n of ψ and ⟨·, ·⟩Sn )
n! n!
w∈Sn w∈Sn
X |Cλ | X |Cλ |
= pλ = ελ pλ
n! n!
λ⊢n λ⊢n
X pλ X pλ
= = ελ
zλ zλ
λ⊢n λ⊢n
= hn = en (by Corollary 9.9.3).
238
Now
ℓ ℓ ℓ
!
Y Y Y
hλ = hλi = ch(χλtriv
i
) = ch χλtriv
i
= ch(IndS n
Sλ χtriv )
n
(7), (8): Each of (4) and (5) says that ch−1 (ΛZ ) is contained in the space of virtual characters, because {hλ }
and {eλ } are Z-module bases for ΛZ , and their inverse images under ch are genuine characters. On the
other hand, {sλ } is also a Z-basis, so each σλ := ch−1 (sλ ) is a character. Moreover, since ch is an isometry
we have
⟨σλ , σµ ⟩Sn = ⟨sλ , sµ ⟩Λ = δλµ
which must mean that {σλ : λ ⊢ n} is a Z-basis for CℓV (Sn ), and that each σλ is either an irreducible
character or its negative. Thus, up to sign changes and permutations, the class functions σλ are just the
characters χλ of the Specht modules indexed by λ (see §8.10). That is, σλ = ±χπ(λ) , where π is a permutation
of Par preserving size.
In fact, we claim that σλ = χλ for all λ. First, we confirm that the signs are positive. We can write each
Schur function as X pµ
sλ = bλ,µ (9.27)
zµ
µ⊢n
so that bλ,µ = ±χπ(λ) (Cµ ). In particular, taking µ = (1n ), the cycle-shape of the identity permutation, we
have
bλ,(1n ) = ± dim χπ(λ) . (9.29)
On the other hand, the only power-sum symmetric function that contains the squarefree monomial x1 x2 · · · xn
is p(1n ) (with coefficient z(1n ) = n!). Extracting the coefficients of that monomial on both sides of (9.27) gives
f λ = bλ,µ . (9.30)
In particular, comparing (9.29) and (9.30), we see that the sign ± is positive for every λ. (We also have a
strong hint that π is the identity permutation, because dim χπ(λ) = f λ .)
239
Proof. We calculate the multiplicity of each irrep in the tabloid representation using characters:
The Frobenius characteristic allows us to translate back and forth between symmetric functions and charac-
ters of symmetric groups. In particular, many questions about representations of Sn can now be answered
in terms of tableau combinatorics. Here are a few fundamental things we would like to know at this point.
1. Irreducible characters. What is the value of the irreducible character χλ = ch−1 (sλ ) on the conjugacy
class Cµ ? In other words, what is the character table of Sn ? We have worked out some examples (e.g., n = 3,
n = 4) and know that the values are all integers, since the Schur functions are an integral basis for Λn . A
precise combinatorial formula is given by the Murnaghan–Nakayama Rule. (According to Stanley [Sta99,
p.410], this formula was first published by Littlewood and Richardson in 1934, predating Murnaghan (1937)
and Nakayama (1941).)
2. Dimensions of irreducible characters. A special case of the Murnaghan–Nakayama Rule is that the
irreducible representation with character χλ has dimension f λ , the number of standard tableaux of shape λ.
What are the numbers f λ ? There is a beautiful interpretation called the hook-length formula of Frame,
Robinson and Thrall, which again has many, many proofs in the literature.
3. Littlewood–Richardson numbers. Now that we know how important the Schur functions are from a
representation-theoretic standpoint, how do we multiply them? That is, suppose that µ, ν are partitions
with |µ| = q, |ν| = r. Then sµ sν ∈ Λq+r , so it has a unique expansion as a linear combination of Schur
functions: X
sµ sν = cλµ,ν sλ , cλµ,ν ∈ Z. (9.31)
λ
The cλµ,ν ∈ Z are called the Littlewood–Richardson numbers. They are the structure coefficients for Λ, re-
garded as an algebra generated as a vector space by the Schur functions. The cλµ,ν must be integers, because
sµ sν is certainly a Z-linear combination of the monomial symmetric functions, and the Schur functions are
a Z-basis.
cλµ,ν = IndS
Sq ×Sr (χµ ⊗ χν ), χλ
n
Sn
= χµ ⊗ χν , ResS
Sq ×Sr (χλ )
n
Sq ×Sr
240
Any combinatorial interpretation for the numbers cλµν is called a Littlewood–Richardson rule; there are
many of them.
4. Transition matrices. What are the coefficients of the transition matrices between different bases of Λn ?
We have worked out a few cases using the Cauchy kernel, and we have defined the Kostka numbers to be
the transition coefficients from the m’s to the s’s (this is just the definition of the Schur functions).
For this section, we will regard symmetric functions as polynomials rather than power series, for a reason
that will quickly become apparent.
Definition 9.13.1. Let Sn act on R = k[x1 , . . . , xn ] by permuting variables. A polynomial a ∈ R is alternat-
ing, or an alternant, if w(a) = ε(w)a for all w ∈ Sn . Equivalently, interchanging any two variables xi , xj
maps a to −a.
In particular, every alternant is divisible by xj − xi for each i < j, hence by the Vandermonde determinant
xn−1
1 xn−2
1 ··· x1 1
Y xn−1
2 xn−2
2 ··· x2 1
V = (xi − xj ) = .. .. .. .. .
1≤i<j≤n . . . .
xn−1
n xn−2
n ··· xn 1
(Why does this equality hold? Interchanging xi with xj swaps two rows of the determinant, hence changes
its sign. Therefore the determinant is divisible by the product on the left. On the other hand, both polyno-
mials are homogeneous of degree n2 = 0 + 1 + · · · + (n − 1), and the coefficients of xn−1
1 xn−2
2 · · · x1n−1 x0n are
both +1, so equality must hold.) This is why we are working with polynomials: replacing 1 ≤ i < j ≤ n
with 1 ≤ i < j ≤ ∞ in the definition of V does not result in a well-defined power series.
We can construct more general alternants by changing the powers of variables that occur in each column of
the Vandermonde determinant: for α = (α1 , . . . , αn ) ∈ Nn , we define
α n
X
aα = aα (x1 , . . . , xn ) = xi j i,j=1 = ε(w)w(xα ). (9.32)
w∈Sn
Note that aα = 0 if (and only if) α contains some entry more than once. Moreover, the entries of α might as
well be listed in decreasing order. Therefore we can write α = λ + δ, where λ = (λ1 ≥ · · · ≥ λn ≥ 0) ∈ Par
and δ = (n − 1, n − 2, . . . , 1, 0), and addition is componentwise: αj = λj + δj = λj + n − j. That is,
n
λ +n−j
aλ+δ = xi j .
i,j=1
In particular, aδ = V . As observed above, every alternant is divisible by aδ , so the quotient aλ+δ /aδ is a
polynomial; moreover, it is a symmetric polynomial, since each w ∈ Sn scales it by ε(w)/ε(w) = 1.
Theorem 9.13.2. For all λ, we have aλ+δ /aδ = sλ (x1 , . . . , xn ).
241
Proof. In light of the second assertion of Corollary 9.10.13 (specialized to the first n variables) and the in-
vertibility of the matrix [Kλµ ], it is equivalent to show that for every µ = (µ1 , . . . , µk ) we have
X
eµ = Kλ̃µ aλ+δ /aδ
λ
or equivalently X
aδ eµ = Kλ̃µ aλ+δ .
λ
Both sides of the equation are alternating, so it is enough to show that for every λ, the monomial xλ+δ
has the same coefficient on both sides of this equation. On the RHS this coefficient is Kλ̃µ since the mono-
mial only appears in the λ summand. On the LHS, the coefficient [xλ+δ ]aδ eµ is the sum of ε(w) over all
factorizations 1 k 1 k
n−1
xλ+δ = w(xδ ) xβ · · · xβ = x0w(1) x1w(2) · · · xw(n) xβ · · · xβ .
i
where each xβ is a squarefree monomial of degree µi . Denote such a factorization by f (w, β) = f (w, β 1 , . . . , β k ),
and denote by F the set of all such factorizations. Thus we are trying to prove that
X
ε(w) = Kλ̃µ . (9.33)
f (w,β)∈F
1 j
Let f (w, β)j denote the partial product w(xδ )xβ · · · xβ . For a monomial M , let powxi (M ) denote the
power of xi that appears in M .
We now describe a sign-reversing involution on the set F . Suppose that f (w, β) is a factorization such that
for some j ∈ [k] and some a ̸= b
For example, let n = 3, λ = (2, 2, 1), α = (4, 3, 1), µ = (2, 2, 1). The set F contains eight factorizations of
xα = x41 x32 x3 , including three cancelling pairs:
1 2 3
w ε(w) w(xδ ) xβ xβ xβ j, {a, b}
123 1 x21 x2 x1 x2 x1 x2 x3 −
123 1 x21 x2 x1 x2 x1 x3 x2 −
123 1 x21 x2 x1 x3 x1 x2 x2 1, {2, 3}
132 −1 x21 x3 x1 x2 x1 x2 x2 1, {2, 3}
123 1 x21 x2 x1 x2 x2 x3 x1 2, {1, 2}
213 −1 x22 x1 x1 x2 x1 x3 x1 2, {1, 2}
123 1 x21 x2 x2 x3 x1 x2 x1 1, {1, 2}
213 −1 x22 x1 x1 x3 x1 x2 x1 1, {1, 2}
The uncanceled factorizations f (w, β) are those for which, in every partial product f (w, β)j all variables
occur with different powers. But in fact this condition implies w = Id, for otherwise, there are indices a < b
for which
powa (w(xδ )) = powa (f (w, β)0 ) < powb (f (w, β)0 ) = powb (w(xδ )) but certainly
δ+λ δ+λ
powa (x ) = powa (f (w, β)k ) > powb (f (w, β)k ) = powb (x )
242
i
but since the xβ are all squarefree, there must be some j such that
In particular, the coefficient [xλ+δ ]aδ eµ is positive: it is the number of factorizations of xλ into squarefree
1 k
monomials xβ , . . . , xβ of degrees µ1 , . . . , µk so that for all j ≤ k we have
1 j 1 j 1 j
pow1 (xβ · · · xβ ) ≥ pow2 (xβ · · · xβ ) ≥ · · · ≥ pown (xβ · · · xβ ). (9.34)
i
Thus each variable xj must occur in λi of the monomials xβ . We record the list of monomials by a tableau
i
of content µ whose entries correspond to monomials xβ and whose columns correspond to variables xj :
column j contains an i if xj occurs in xαi . Thus the tableau has shape λ̃. We can arrange each column
in increasing order, so the the entry in (i, j) tells us the ith monomial divisible by xj . Continuing our
example, the two factorizations of xλ that remain uncancelled (see the preceding table) give rise to tableaux
as follows:
x1 x2 · x1 x2 · x3 x1 x2 · x1 x3 · x2
1 1 3 1 1 2
2 2 2 3
i
There are no repeats in columns because no variable occurs more than once in any xβ . Moreover, if the
ith row has a strict decrease a > b between the jth and (j + 1)st columns, then this means that the ith
occurrence of xj occurs later than the ith occurrence of xi — i.e., there are more xj+1 ’s then xj in the first
b monomials, which contradicts (9.34). Hence the tableau is column-strict. Moreover, every column-strict
tableau of shape λ̃ and content µ gives rise to a factorization that contributes 1 to the coefficient [xλ+δ ]aδ eµ .
We conclude that the coefficient is Kλ̃µ as desired.
We know from Theorem 9.11.2 that the irreducible characters of Sn are χλ = ch−1 (sλ ) for λ ⊢ n. We want to
compute these numbers. Via the Frobenius characteristic, this problem is equivalent to expanding the Schur
functions (which correspond to irreducible characters) as linear combinations of the power-sums (which
correspond to indicator functions of conjugacy classes). We will need the description of Schur functions as
quotients of alternants in §9.13, and the key step will be expressing a product sν pr as a linear combination
of Schur functions (equation (9.36)).
We first state the result, then prove it. The relevant combinatorial objects are ribbons and ribbon tableau.
A ribbon is a connected8 skew shape R with no 2 × 2 block, or equivalently with no square both north and
west of another square. The size |R| is as usual the number of squares in the ribbon, and its height h(R) is
the number of rows.9
8 “Connected” means “connected with respect to sharing edges, not just diagonals”, or equivalently “the topological interior is
connected”; for example, the skew shape 21/1 = is not considered to be connected.
9 Stanley defines the height as one less than the number of rows, which simplifies the formulas but seems less natural to me.
243
A ribbon tableau is a decomposition of a Ferrers diagram into ribbons R1 , . . . , Rk such that for each i ≤ k,
the union of the first i ribbons forms a Ferrer diagram. Here is an example of a ribbon tableau of shape
λ = (8, 7, 6, 6, 4) into k = 6 ribbons.
1 1 1 3 4 4 4 4
1 2 3 3 4 6 6
1 2 3 4 4 6
1 2 5 6 6 6
5 5 5 6
Note that each row and column is weakly increasing, and that for each i ≤ k, the union R1 ∪ · · · ∪ Ri is a
partition. In this context ribbons are often called border strips or rim hooks.
The list ρ of sizes of the ribbons is the content of the ribbon tableau; here ρ = (6, 3, 4, 7, 4, 7). Let RT (λ, ρ)
denote the set of ribbon tableaux of shape λ and content ρ, and for T = (R1 , . . . , Rk ) ∈ RT (λ, ρ) put
k
Y
(−1)T = (−1)1+ht(Ri ) .
i=1
For example, the heights of R1 , . . . , R6 in the ribbon tableau T shown above are 4,3,3,3,2,4. There are an
odd number of even heights, so (−1)T = −1.
1. λi ∈ [νi + 1, νi−1 ].
2. For each k ∈ [i + 1, j] we have λk = νk−1 + 1.
Proof. For (1), we have λi ≤ λi−1 = νi−1 ; on the other hand, λi is obtained by adding at least one box to
νi . (In particular this interval cannot be empty — it is possible to add at least one box in the ith row of ν
without changing the (i − 1)st row, so it must be the case that νi−1 > νi .)
(2) asserts that the last box in the kth row of λ must be one column east and one column south of the last
box in the (k − 1)st row of ν. Indeed, any further west and R would not be connected; any further east and
it would have a 2 × 2 block.
αi = νi + n − i.
244
Let aα be the alternant of (9.32), and let ϵj be the sequence with a 1 in position j and 0s elsewhere. For
r ∈ N, we have
X
aα pr (x1 , . . . , xn ) = ε(w)w(xα )(xr1 + · · · + xrn )
w∈Sn
X
= ε(w)xα1 αn r r
w(1) · · · xw(n) (xw(1) + · · · + xw(n) )
w∈Sn
X n
X
= ε(w) w(xα+rϵj )
w∈Sn j=1
n
X X
= ε(w)w(xα+rϵj )
j=1 w∈Sn
n
X
= aα+rϵj . (9.35)
j=1
If two entries of α + rϵj are equal, then aα+rϵj = 0. Otherwise, there is some i ∈ [j] such that
or equivalently
νi−1 + n − (i − 1) > νj + n − j + r > νi + n − i.
(If i = 1, just ignore the first inequality.) Therefore, sorting the parts of α + rϵj in decreasing order means
moving the j th part back to position i and pushing parts i, i + 1, . . . , j − 1 up — that is, acting by a (j − i + 1)-
cycle, which has sign (−1)j−i . That is, aα+rϵj = (−1)j−i aλ+δ , where
Now a miracle occurs: by Lemma 9.14.1, these partitions λ are precisely the ones for which λ/ν is a ribbon
of size r, spanning rows i, . . . , j and hence of height j − i + 1. Combining this observation with (9.35) we
get
n
X
aν+δ pr = aα pr = aα+rϵj
j=1
X
= (−1)ht(R)+1 aλ+δ
R,λ
where the sum runs over ribbons R of size r that can be added to ν to obtain a partition λ. Dividing both
sides by aδ and applying Theorem 9.13.2 gives
X
sν pr = (−1)ht(R)+1 sλ . (9.36)
R,λ
(This is valid on the level of power series as well as for polynomials, since it remains valid under increasing
the number of variables, so the coefficient of every monomial in the power series is equal on both sides.)
X k
Y
sν pµ = (−1)ht(Ri )+1 sλ (9.37)
R1 ,...,Rk ,λ i=1
245
where the sum runs over k-tuples of ribbons of lengths given by the parts of µ that can be added to ν to
obtain λ. In particular, if ν = ∅, then this is simply the statement that T = (R1 , . . . , Rk ) is a ribbon tableau
of shape λ and content µ, and the sign is (−1)T , so we get
X X
pµ = (−1)T sλ (9.38)
λ T ∈RT (λ,µ)
so that
(−1)T = ⟨pµ , sλ ⟩Λ = ⟨ch−1 (pµ ), ch−1 (sλ )⟩Sn (since ch−1 is an isometry)
X
T ∈RT (λ,µ)
As a first consequence, we can expand the Schur functions in the power-sum basis:
Corollary 9.14.3. For all λ ⊢ n we have
X pµ X εµ p µ
sλ = χλ (Cµ ) and sλ̃ = χλ (Cµ ) .
µ
zµ µ
zµ
P
Proof. Write sλ in the p-basis as µ bλµ pµ . Taking the Hall inner product of both sides with pµ gives
⟨sλ , pµ ⟩ = bλµ zµ , or bλµ = zµ−1 ⟨sλ , pµ ⟩, implying the first equality. Applying ω and invoking Corollar-
ies 9.8.2 and 9.9.33 gives the second equality.
An important special case of the Murnaghan–Nakayama rule is when µ = (1, 1, . . . , 1), since then χλ (Cµ ) =
χλ (IdSn ), is just the dimension of the irreducible character χλ . On the other hand, a ribbon tableau of
content µ is just a standard tableau. So the Murnaghan–Nakayama Rule implies the following:
Corollary 9.14.4. dim χλ = f λ , the number of standard tableaux of shape λ.
246
T
P P
We calculate λ T ∈RT (λ,ρ) (−1) for each permutation ρ of µ:
1 2 2
(−1)1+0+0 = −1
1 3
1 2 3 (−1)1+1+0 = 1
1 2
Let λ ⊢ n, let ℓ = ℓ(λ), and let SYT(λ) the set of standard tableaux of shape λ, so f λ = |SYT(λ)|. In what
follows, we label the rows and columns of a tableau starting at 1. If c = (i, j) is the cell in the ith row and
jth column of a tableau T , then T (c) or T (i, j) denotes the entry in that cell.
The hook H(c) defined by a cell c = (i, j) consists of itself together with all the cells due east or due south
of it. The number of cells in the hook is the hook length, written h(c) or h(i, j). (In this section, the letter h
always refers to hook lengths, never to the complete homogeneous symmetric function.) In the following
example, h(c) = h(2, 3) = 6.
1 2 3 4 5
1
2 c
3
4
5
6
247
Theorem 9.15.1 (Hook-Length Formula). Let λ ⊢ n. Then the number f λ of standard Young tableaux of shape λ
equals F (λ), where
n!
F (λ) = Y .
h(c)
c∈λ
9 7 6 3 1
7 5 4 1
5 3 2
4 2 1
1
Before getting started, here is how not to prove the hook-length formula. Consider the discrete probability
space of all n! fillings of the Ferrers diagram of λ with the numbers 1, . . . , n. Let S be the event that a
uniformly chosen filling T is a standard tableau,
T and for each cell, let Xc be the event that T (c) is the
number in the hook H(c). Then S = c Xc , and Pr[Xc ] = 1/h(c). We would like to conclude that
smallest Q
Pr[S] = c 1/h(c), which would imply the hook-length formula. However, that inference would require
that the events Xc are mutually independent, which they certainly are not! Still, this is a nice heuristic
argument (attributed by Wikipedia to Knuth) that one can at least remember.
There are many proofs of the hook-length formula in the literature. This one is due to Greene, Nijenhuis
and Wilf [GNW79].
Proof of Theorem 9.15.1. First, observe that for every T ∈ SYT(λ), the cell c ∈ T containing the number
n = |λ| must be a corner of λ (i.e., the rightmost cell in its row and the bottom cell in its column). Deleting c
produces a standard tableau of size n − 1; we will call the resulting partition λ − c. This construction gives
a collection of bijections
{T ∈ SYT(λ) : T (c) = n} → SYT(λ − c)
for each corner c.
Now to the main argument. We will prove by induction on n that f λ = F (λ). The base case n = 1 is clear.
For the inductive step, we wish to show that
X X F (λ − c)
F (λ) = F (λ − c) or equivalently =1 (9.39)
corners c corners c
F (λ)
since by the inductive hypothesis together with the bijections just described, the right-hand side of this
equation equals f λ .
Let c = (x, y) be a corner cell. Removing c decreases by 1 the sizes of the hooks H(c′ ) for cells c′ strictly
248
north or west of c, and leaves all other hook sizes unchanged. Therefore,
x−1 y−1
F (λ − c) (n − 1)! Y h(i, y) Y h(x, j)
=
F (λ) n! i=1
h(i, y) − 1 j=1 h(x, j) − 1
x−1 y−1
Y
1 Y 1 1
= 1+ 1+
n i=1 h(i, y) − 1 j=1 h(x, j) − 1
!
1 X Y 1 Y 1
= . (9.40)
n h(i, y) − 1 h(x, j) − 1
A⊆[x−1] i∈A j∈B
B⊆[y−1]
Consider the following random process (called a hook walk). First choose a cell (a0 , b0 ) uniformly from λ.
Then for each t = 1, 2, . . . , move to a cell (at , bt ) chosen uniformly from all other cells in H(at−1 , bt−1 ). The
process stops when it reaches a corner; let pc be the probability of reaching a particular corner c. Evidently
P
c pc = 1. Our goal now becomes to show that
F (λ − c)
pc = (9.41)
F (λ)
which will establish (9.39).
Consider a hook walk starting at (a, b) = (a1 , b1 ) and ending at (am , bm ) = (x, y). Let A = {a1 , . . . , am } and
B = {b1 , . . . , bm } be the sets of rows and columns encountered (removing duplicates); call these sets the
horizontal and vertical projections of W . Let
p(A, B a, b)
denote the probability that a hook walk starting at (a, b) has projections A and B. We claim that
Y 1 Y 1
p(A, B a, b) = . (9.42)
h(i, y) − 1 h(x, j) − 1
i∈A\x j∈B\y
| {z }
Φ
We prove this by induction on m. If m = 1, then either A = {a} = {x} and B = {b} = {y}, and the equation
reduces to 1 = 1 (the RHS is the empty product), or else it reduces to 0 = 0. If m > 1, then
p(A \ a1 , B a2 , b1 ) p(A, B \ b1 a1 , b2 )
p(A, B a, b) = +
h(a, b) − 1 h(a, b) − 1
| {z } | {z }
first move south to (a2 , b1 ) first move east to (a1 , b2 )
1
= (h(a, y) − 1)Φ + (h(x, b) − 1)Φ (by induction)
h(a, b) − 1
h(a, y) − 1 + h(x, b) − 1
= Φ. (9.43)
h(a, b) − 1
To see that the parenthesized expression in (9.43) is 1, consider the following diagram, with the hooks at
(a, y) and (x, b) shaded in red and blue respectively, with the corner (x, y) omitted so that there are a total of
h(a, y) − 1 + h(x, b) − 1 shaded cells. Pushing some red cells north and some blue cells to the left produces
the hook at (a, b) with one cell omitted, as on the right.
249
(a, b)
(x, b)
(a, y)
(x, y)
This proves (9.42). Now we compute pc , the probability that a walk ends at a particular corner c = (x, y).
Equivalently, x ∈ A and y ∈ B; equivalently, A ⊆ [x] and B ⊆ [y]. Therefore, summing over all possible
starting positions, we have
1 X
pc = p(A, B a, b)
n
(A,B,a,b):
A⊆[x], B⊆[y]
a=min A, b=min B
x=max A, y=max B
1 X Y 1 Y 1
= (by (9.42))
n h(i, y) − 1 h(x, j) − 1
(A,B,a,b) i∈A\x j∈B\y
as above
!
1 X Y 1 Y 1
=
n h(i, y) − 1 h(x, j) − 1
A⊆[x−1] i∈A j∈B
B⊆[y−1]
which is precisely (9.40). This establishes (9.41) and completes the proof.
UNDER CONSTRUCTION
Recall that the Littlewood–Richardson coefficients cλµν are the structure coefficients for Λ as an algebra with
vector space basis {sλ : λ ∈ Par}: that is,
X
sµ sν = cλµν sλ .
λ
where c̃λ/µ,ν ∈ Z for all λ, µ, ν. In fact these numbers are also Littlewood–Richardson coefficients, and they
are symmetric in µ and ν (which is hardly obvious from the definition).
250
Proposition 9.16.1. Let x = {x1 , x2 , . . . }, y = {y1 , y2 , . . . } be two countably infinite sets of variables. Then
X
sλ (x, y) = sµ (x)sλ/µ (y).
µ⊆λ
Proof. Consider column-strict tableaux of shape λ with labels taken from the alphabet 1 < 2 < · · · < 1′ <
2′ < · · · , and let the weight of such a tableau T be xα yβ , where αi (resp., βi ) is the number of cells filled
with i (resp., i′ ). Then the left-hand side is the generating function for all schools tableaux by weight. On
the other hand, such a tableau consists of a CST of shape µ filled with 1, 2, . . . (for some µ ⊆ λ) together
with a CST of shape λ/µ filled with 1′ , 2′ , . . . , so the RHS enumerates the same set of tableaux.
Theorem 9.16.2. For all partitions λ, µ, ν, we have
Equivalently,
⟨sµ sν , sλ ⟩Λ = ⟨sν , sλ/µ ⟩Λ .
Proof. We need three countably infinite sets of variables x, y, z for this. Consider the “double Cauchy ker-
nel” Y Y
Ω(x, z)Ω(y, z) = (1 − xi zj )−1 (1 − yi zj )−1 .
i,j i,j
On the one hand, expanding both factors in terms of Schur functions and then applying the definition of
the Littlewood–Richardson coefficients to the z terms gives
! !
X X X
Ω(x, z)Ω(y, z) = sµ (x)sµ (z) sν (y)sν (z) = sµ (x)sν (y)sµ (z)sν (z)
µ ν µ,ν
X X
= sµ (x)sν (y) cλµ,ν sλ (z). (9.44)
µ,ν λ
(The first equality is perhaps clearer in reverse; think about how to express the right-hand side as an infinite
product over the variable sets x ∪ y and z. The second equality uses Proposition 9.16.1.) Now the theorem
follows from the equality of (9.44) and (9.45).
There are a lot of combinatorial interpretations of the Littlewood–Richardson numbers. Here is one. A
ballot sequence (or Yamanouchi word, or lattice permutation) is a sequence of positive integers such that
each initial subsequence contains at least as many 1’s as 2’s, at least as many 2’s as 3’s, et cetera.
Theorem 9.16.3 (Littlewood–Richardson Rule). cλµ,ν equals the number of column-strict tableaux T of shape λ/µ,
and content ν such that the word obtained by reading the entries of T row by row, right to left, top to bottom, is a
ballot sequence.
251
Include a proof. There are a lot of them but they tend to be hard.
Important special cases are the Pieri rules, which describe how to multiply by the Schur function corre-
sponding to a single row or column (i.e., by an h or an e.)
Theorem 9.16.4 (Pieri Rules). Let (k) denote the partition with a single row of length k, and let (1k ) denote the
partition with a single column of length k. Then
X
sµ s(k) = sµ hk = sλ
λ
where λ ranges over all partitions obtained from µ by adding k boxes, no more than one in each column; and
X
sµ s(1k ) = sµ ek = sλ
λ
where λ ranges over all partitions obtained from µ by adding k boxes, no more than one in each row.
where λ ranges over all partitions obtained from µ by adding a single box. Via the Frobenius characteristic,
this gives a “branching rule” for how the restriction of an irreducible character of Sn splits into a sum of
irreducibles when restricted:
ResSSn−1 (χλ ) = ⊕µ χµ
n
where now µ ranges over all partitions obtained from λ by deleting a single box. Details?
Definition 9.17.1. Let b, b′ be finite ordered lists of positive integers (or “words in the alphabet N>0 ”). We
say that b, b′ are Knuth equivalent, written b ∼ b′ , if one can be obtained from the other by a sequence of
transpositions as follows:
(Here the notation · · · xzy · · · means a word that contains the letters x, z, y consecutively.)
For example, 21221312 ∼ 21223112 by Rule 1, and 21223112 ∼ 21221312 by Rule 2 (applied in reverse).
This definition looks completely unmotivated at first, but hold that thought!
We now define an equivalence relation on column-strict skew tableaux, called jeu de taquin10 . The rule is
as follows:
• y x≤y
−−−→ x y • y x>y
−−−→ y •
x • x x
10 French for “sliding game”, roughly; it refers to the 15-square puzzle with sliding tiles that used to come standard on every
252
That is, for each inner corner of T — that is, an empty cell that has numbers to the south and east, say x
and y — then we can either slide x north into the empty cell (if x ≤ y) or slide y west into the empty cell
(if x > y). It is not hard to see that any such slide (hence, any sequence of slides) preserves the property of
column-strictness.
For example, the following is a sequence of jeu de taquin moves. The bullets • denote the inner corner that
is being slid into.
• 1 4 → 1 1 4 → 1 1 4 → 1 1 4 → 1 1 4
1 2 • 2 2 • • 2 4 2 2 4
2 3 4 2 3 4 2 3 4 2 3 • 3
(9.46)
→ • 1 1 4 → 1 • 1 4 → 1 1 • 4 → 1 1 4 4
2 2 4 2 2 4 2 2 4 2 2
3 3 3 3
If two skew tableaux T, T ′ can be obtained from each other by such slides (or by their reverses), we say that
they are jeu de taquin equivalent, denoted T ≈ T ′ . Note that any skew column-strict tableau T is jeu de
taquin equivalent to an ordinary CST (called the rectification of T ); see, e.g., the example (9.46) above. In
fact, the rectification is unique; the order in which we choose inner corners does not matter.
Definition 9.17.2. Let T be a column-strict skew tableau. The row-reading word of T , denoted row(T ), is
obtained by reading the rows left to right, bottom to top.
For example, the reading words of the skew tableaux in (9.46) are
If T is an ordinary (not skew) tableau, then it is determined by its row-reading word, since the “line breaks”
occur exactly at the strict decreases of row(T ). For skew tableaux, this is not the case. Note that some of
the slides in (9.46) do not change the row reading word; as a simpler example, the following skew tableaux
both have reading word 122:
1 2 2 2 2 2
1 2 1
On the other hand, it’s not hard to se that rectifying the second or third tableau will yield the first; therefore,
they are all jeu de taquin equivalent.
For a word b on the alphabet N>0 , let P(b) denote its insertion tableau under the RSK algorithm. (That is,
construct a generalized permutation bq in which q is any word; run RSK; and remember only the tableau
P , so that the choice of q does not matter.)
Theorem 9.17.3. (Knuth–Schützenberger) For two words b, b′ , the following are equivalent:
1. P (b) = P (b′ ).
2. b ∼ b′ .
3. T ≈ T ′ , for any (or all) column-strict skew tableaux T, T ′ with row-reading words b, b′ respectively.
This is sometimes referred to (e.g., in [Ful97]) as the equivalence of “bumping” (the RSK algorithm as
presented in Section 9.10) and “sliding” (jeu de taquin).
253
9.18 Yet another version of RSK
Fix w ∈ Sn . Start by drawing an n × n grid, numbering columns west to east and rows south to north. For
each i, place an X in the i-th column and wi -th row. We are now going to label each of the (n + 1) × (n + 1)
intersections of the grid lines with a partition, such that the partitions either stay the same or get bigger as
we move north and east. We start by labeling each intersection on the west and south sides with the empty
partition ∅.
8 ×
7 ×
6 ×
5 ×
4 ×
3 ×
2 ×
1 ×
1 2 3 4 5 6 7 8
For each box whose SW, SE and NW corners have been labeled λ, µ, ν respectively, label the NE corner ρ
according to the following rules:
Rule 2: If λ ⊊ µ = ν and the box doesn’t contain an X, then it must be the case that µi = λi + 1 for some i.
Obtain ρ from µ by incrementing µi+1 .
Rule 3: If µ ̸= ν, then set ρ = µ ∨ ν (where ∨ means the join in Young’s lattice: i.e., take the componentwise
maximum of the elements of µ and ν).
Rule X: If there is an X in the box, then it must be the case that λ = µ = ν. Obtain ρ from λ by incrementing
λ1 .
Note that the underlined assertions need to be proved; this can be done by induction.
Example 9.18.1. Let n = 8 and w = 57214836. In Example 9.10.2, we found that RSK(w) = (P, Q), where
P = 1 3 6 and Q= 1 2 6 .
2 4 8 3 5 8
5 7 4 7
The following extremely impressive figure shows what happens when we run the alternate RSK algorithm
on w. The partitions λ are shown in red. The numbers in parentheses indicate which rules were used.
254
0 1 2 21 211 221 321 322 332
(3) (3) (3) (3) (3) (3) (2)
0 1 2 21 211 221 221 222 322
(3) (3) (3) (2) (3) (2) (3)
0 1 1 11 111 211 211 221 321
(3) (1) (3) (3) (3) (1) (3)
0 1 1 11 111 211 211 221 221
(3) (2) (2) (3) (3) (3) (3)
0 0 0 1 11 21 21 22 22
(1) (1) (3) (3) (3) (2) (3)
0 0 0 1 11 11 11 21 21
(1) (1) (3) (3) (1) (1) (3)
0 0 0 1 11 11 11 11 11
(1) (1) (2) (3) (3) (3) (3)
0 0 0 0 1 1 1 1 1
(1) (1) (1) (3) (3) (3) (3)
0 0 0 0 0 0 0 0 0
Observe that:
• Rule 1 is used exactly in those squares that have no X either due west or due south.
• For all squares s, |ρ| is the number of X’s in the rectangle whose northeast corner is s. In particular,
the easternmost partition λ(k) in the kth row, and the northernmost partition µ(k) in the kth column,
both have size k.
• It follows that the sequences
∅ = λ(0) ⊆ λ(1) ⊆ · · · ⊆ λ(n) ,
∅ = µ(0) ⊆ µ(1) ⊆ · · · ⊆ µ(n)
correspond to SYT’s of the same shape (in this case 332).
• These SYT’s are the P and Q of the RSK correspondence!
Definition 9.19.1. A quasisymmetric function is a formal power series F ∈ C[[x1 , x2 , . . . ]] with the fol-
lowing property: if i1 < · · · < ir and j1 < · · · < jr are two sets of indices in strictly increasing order and
α1 , . . . , αr ∈ N, then
[xα αr α1 αr
i1 · · · xir ]F = [xj1 · · · xjr ]F
1
255
Symmetric functions are automatically quasisymmetric, but not vice versa. For example,
X
x2i xj
i<j
is quasisymmetric but not symmetric (in fact, it is not preserved by any permutation of the variables). On
the other hand, the set of quasisymmetric functions forms a graded ring QSym ⊆ C[[x]]. We now describe
a vector space basis for QSym.
A composition α is a sequence (α1 , . . . , αr ) of positive integers, called its parts. Unlike a partition, we do
not require that the parts be in weakly decreasing order. If α1 + · · · + αr = n, we write α |= n; the set
of all compositions of n will be denoted Comp(n). Sorting the parts of a composition in decreasing order
produces a partition of n, denoted by λ(α).
Compositions are much easier to count than partitions. Consider the set of partial sums
S(α) = {α1 , α1 + α2 , . . . , α1 + · · · + αr−1 }.
The map α 7→ S(α) is a bijection from compositions of n to subsets of [n−1]; in particular, | Comp(n)| = 2n−1 .
We can define a partial order on Comp(n) via S by setting α ⪯ β if S(α) ⊆ S(β); this is called refinement.
The covering relations are merging two adjacent parts into one part.
i1 <···<ir
Just as for the monomial symmetric functions, every monomial appears in exactly one Mα , and Defini-
tion 9.19.1 says precisely that a power series f is quasisymmetric if all monomials appearing in the same
Mα have the same coefficient in f . Therefore, the set {Mα } is a graded basis for QSym.
Example 9.19.2. Let M be a matroid on ground set E of size n. Consider weight functions f : E → N>0 ;
one of the definitions of a matroid (see the problem set) is that a smallest-weight basis of M can be chosen
via the following greedy algorithm (list E in weakly increasing order by weight e1 , . . . , en ; initialize B = ∅;
for i = 1, . . . , n, if B + ei is independent, then replace B with B + ei ). The Billera-Jia-Reiner invariant of M
is the formal power series X
W (M) = xf (1) xf (2) · · · xf (n)
f
where the sum runs over all weight functions f for which there is a unique smallest-weight basis. The
correctness of the greedy algorithm implies that W (M) is quasisymmetric.
For example, let E = {e1 , e2 , e3 } and M = U2 (3). The bases are e1 e2 , e1 e3 , and e2 e3 . Then E has a unique
smallest-weight basis iff f has a unique maximum; it doesn’t matter if the two smaller weights are equal or
not. If the weights are all distinct then they can be assigned to E in 3! = 6 ways; if the two smaller weights
are equal then there are three choices for the heaviest element of E. Thus
X X
W (U2 (3)) = 6xi xj xk + 3xi x2j = 6M111 + 3M12 .
i<j<k i<j
256
9.20 Exercises
Problem 9.1. Suppose λ ⊢ n and µ ⊢ m are partitions. Then the product mλ mµ is a symmetric function of
degree m + n, so there is a unique expression
X
mλ mµ = aλ,µ
ν mν
ν⊢m+n
Solution: Let aλ,µ be the coefficient of hλ when expanded in the monomial basis, that is,
X
hλ = aλ,µ mµ .
µ
By analogy with the proof of Theorem 9.4.2, the number aλ,µ is the number of λ-factorizations of xµ into
arbitrary (not necessarily squarefree) monomials xα1 , . . . , xαℓ , and such factorizations are in bijection with
tableaux of shape λ and content µ, whose rows are weakly (as opposed to strictly) increasing. Now, by
analogy with the proof of Corollary 9.4.4, such a tableau also represents a µ-factorization of xλ into arbitrary
monomials xβ1 , . . . , xβk , where each occurrence of j in the ith row corresponds to a factor of xi in xβj .
1 1 2 3
2 2
1 4
1
257
Solution: Write uλ and vµ in the h- and m-bases, respectively:
X X
uλ = aλσ hσ , vµ = bµτ mτ .
σ τ
So for each n, the matrices A = [aλσ ]λ,σ and B = [bµτ ]µ,τ are inverse transposes, since by (9.48) the dot
product of the σth column of A with the τ th column of B is δστ . On the other hand, that means that
multiplying the µth row of B by the λth row of A — which equals ⟨uλ , vµ ⟩ by (9.47) — also gives δλ,µ .
Problem 9.4. More generally, for two graded bases {uλ }, {vµ } of Λ, show how to get the values of ⟨uλ , vµ ⟩Λ
by expanding the Cauchy kernel.
Solution: Throughout, fix n ∈ N and restrict λ, µ, ν, ξ to partitions of n. Write uλ and vλ in the Schur basis:
X X
uλ = aλµ sµ , vλ = bλµ sµ .
µ µ
On the other hand, the piece of Ω that is homogeneous of degree n in each of x, y can be expanded in the
basis {uλ (x)vµ (y)} is X
Ω= cλµ uλ (x)vµ (y)
λ,µ
P
Comparing this with the expansion Ω = ν sν (x)sν (y), we get
X
cλµ aλν bµξ = δν,ξ ∀ν, ξ.
λ,µ
258
This is equivalent to the matrix equation AT CB = Id, where C = [cλµ ]λ,µ⊢n . Equivalently,
−1
C = (AT )−1 B −1 = (BAT )−1 = ((AB T )T )−1 = [⟨uλ , vµ ⟩]µ,λ .
In other words, to find ⟨uλ , vµ ⟩, expand Ω in the basis uλ (x)vλ (y), extract the coefficients, and take the
inverse transpose of the resulting matrix. (Note that this result implies Prop. 9.9.5(a) when the coefficient
matrix is the identity.)
Problem 9.5. Let λ ⊢ n. Verify that |Cλ | = n!/zλ , where zλ is defined as in (9.16).
Solution: Recall the definition of ε from (9.16). The definitions of ε and ℓ make sense for compositions as
well as for partitions. Also, note that the number of compositions whose underlying partition λ ⊢ n (that
is, the number of ways to rearrange the parts of λ into a compositions) is ℓ(λ)!/r1 ! · · · rn !, where ri denotes
the number of i’s in λ (as in (9.16)). We have
m
X 1 X (−1)n+1 xn
exp(1 + x) =
m! n
m≥0 n≥1
X 1 X X (−1)ε(β)
= xn
m! β 1 · · · βm
m≥0 n≥1 β|=n
ℓ(β)=m
X X (−1)ε(β)
= 1+ xn
ℓ(β)! · β1 · · · βℓ(β)
n≥1 β|=n
| {z }
q(n)
We want to show that q(1) = δn,1 (Kronecker delta). Rewriting as a sum over partitions, we obtain
X ℓ(λ)! (−1)ε(λ)
n! · q(n) = n! ·
r1 ! · · · rn ! ℓ(λ)! · λ1 · · · λℓ(λ)
λ⊢n
X n!
= (−1)ε(λ)
r1 ! · · · rn ! · λ1 · · · λℓ(λ)
λ⊢n
X n! X
= (−1)ε(λ) = (−1)ε(λ) |Cλ |
zλ
λ⊢n λ⊢n
= |An | − |Sn \ An | = δn,1 as desired.
259
Problem 9.7. Supply the proofs for the identities (9.18), i.e.,
X X pλ (x)pλ (y)
Ω∗ = eλ (x)mλ (y) = ελ .
zλ
λ λ
X qn
For the second identity, recall the power series expansion log(1 + q) = (−1)n+1 , which gives
n
n≥1
Y X
log (1 + xi yj ) = log(1 + xi yj )
i,j≥1 i,j≥1
X X (−1)n+1 xni yjn X1 X
= = (−1)n+1 xni yjn
n n
i,j≥1 n≥1 n≥1 i,j≥1
X pn (x)pn (y)
= (−1)n+1 .
n
n≥1
X qn
Now exponentiate both sides and apply the power series expansion exp(q) = :
n!
n≥0
X pn (x)pn (y)
Ω∗ = exp (−1)n+1
n
n≥1
k
X 1 X p n (x)p n (y)
= (−1)n+1
k! n
k≥0 n≥1
X 1 k
p1 (x)p1 (y) p2 (x)p2 (y) p3 (x)p3 (y)
= − + − ···
k! 1 2 3
k≥0
r1 (λ) r2 (λ) r3 (λ)
−p2 (x)p2 (y)
X 1
X k p1 (x)p1 (y) p3 (x)p3 (y)
= ···
k! r1 ! r2 ! · · · 1 2 3
k≥0 λ: ℓ(λ)=k
X pλ (x)pλ (y)
= ελ .
zλ
λ
260
Problem 9.8. Prove part (6) of Theorem 9.11.2.
Solution: Recall from Corollary 9.9.33 that ω(pλ ) = ελ pλ . Also, χsign (w) = εsh(w) , since both equal −1 raised
to the power of the number of even cycles in w. Equivalently, on the level of class functions, χsign (Cλ ) = ελ .
Here are two equivalent calculations:
! !
1 X X pλ
ω(ch(χ)) = ω χ(u) psh(u) =ω χ(Cλ )
n! zλ
u∈Sn λ⊢n
1 X X ω(pλ )
= χ(u) ω(psh(u) ) = χ(Cλ )
n! zλ
u∈Sn λ⊢n
1 X X pλ
= χ(u) εsh(u) psh(u) = χ(Cλ ) ελ
n! zλ
u∈Sn λ⊢n
1 X X pλ
= χ(u)χsign (u) psh(u) = χ(Cλ )χsign (Cλ )
n! zλ
u∈Sn λ⊢n
Problem 9.9. Confirm that the Murnaghan–Nakayama rule correctly predicts the values of the trivial, sign,
and standard characters on Sn .
Solution: Fix µ = 1r1 2r2 · · · ⊢ n. If λ = (n) or λ = (1n ) then there is exactly one ribbon tableau of shape λ and
content µ, in which the ribbons are either single-row assembled left to right, or single-column assembled
top to bottom. In the former case all heights are 1 and the sign of the tableau is 1; in the latter case the sign
is −1r2 +r4 +··· , as expected.
1. r1 = n, i.e., µ = 1n . Then ribbon tableaux are just standard tableaux, of which there are n − 1 =
dim χstd , all positive.
2. r1 = 0. Then R1 has to include the cell in the second row, and there is no choice thereafter. So there is
exactly one ribbon tableau, and its sign is negative. Hence χstd (Cµ ) = −1.
3. 1 ≤ r1 < n. If R1 includes the cell in the second row, then there is no choice thereafter, so this way we
get one negative ribbon tableau, as in the previous case. Otherwise, R1 is a single-row shape, when
any part of µ equal to 1 can go in the second row, giving r1 positive ribbon tableau.
In all cases the Murnaghan–Nakayama rule predicts χλ (Cµ ) = r1 − 1, which we know is the value of
χstd (Cµ ).
Problem 9.10. Fill in the proofs of the underlined assertions in Rule 2 and Rule X for the alternate RSK
algorithm in Section 9.18.
Solution: To be written
Problem 9.11. For this problem, you will probably want to use one of the alternate RSK algorithms from
Sections 9.17 and 9.18.
(a) For w ∈ Sn , let (P (w), Q(w)) be the pair of tableaux produced by the RSK algorithm from w. Denote
by w∗ the reversal of w in one-line notation (for instance, if w = 57214836 then w∗ = 63841275). Prove
that P (w∗ ) = P (w)T (where T means transpose).
261
(b) (Open problem) For which permutations does Q(w∗ ) = Q(w)? Computation indicates that the number
of such permutations is
(n−1)/2
2 (n − 1)!
if n is odd,
((n − 1)/2)!2
0 if n is even,
Solution: To be written
Problem 9.12. Let G = (V, E) be a finite simple graph with vertex set V . Let C(G) denote the set of proper
colorings of G: functions κ : V → N>0 such that κ(v) ̸= κ(w) whenever v, w are adjacent in G. Define a
formal power series in indeterminates x1 , x2 , . . . , by
X Y
XG = xκ(v) .
κ∈C(G) v∈V
| {z }
xκ
(a) Show that XG is a symmetric function (this is not too hard). It is known as the chromatic symmetric
function, and was introduced by Stanley [Sta95];
(b) Determine XG for (i) Kn ; (ii) Kn (i.e., the graph with n vertices and no edges); (iii) the four simple
graphs on 3 vertices; (iv) the two trees on 4 vertices.
(c) Explain how to recover the chromatic polynomial pG (k) (see Example 2.3.5) from XG . Does pG (k)
determine XG ?
(d) For a set A ⊆ E, let λ(A) denote the partition whose parts are the sizes of the components of the
subgraph G|A induced by A (so λ ⊢ |V (G)| and ℓ(λ) is the number of components). Prove [Sta95,
Thm. 2.5] that the expansion of XG in the power-sum basis is
X
XG = (−1)|A| pλ(A) .
A⊆E
Solution:
(a) Permuting the variables {xi } corresponds to permuting colors, which preserves the property of being
a proper coloring.
(b) XKn = n!en and XKn = en1 . Let Gj be the simple graph with 3 vertices and j edges; then
262
(c) For a fixed k, setting x1 = · · · = xk = 1 and xj = 0 for all j > k recovers the number of colorings
using only colors 1, . . . , k, i.e., the number pG (k). Thus XG contains all the data of pG (k), although to
actually calculate the latter as a polynomial in k we would need to do more work, e.g., find the values
pG (k) for sufficiently many k and then use Lagrange interpolation. The example of the two degree-4
trees shows that one cannot recover XG from pG (k).
(d) For a coloring κ, let I(κ) denote the set of improper edges of κ. For an edge set A ⊆ E, let
X
f (A) = xκ ,
κ:V (G)→N>0 : I(κ)=A
X
g(A) = xκ
κ:V (G)→N>0 : I(κ)⊇A
P
so that g(A) = B⊇A f (B). Note that a coloring κ with I(κ) ⊇ A is determined by choosing a color
for each component of λ(A). Therefore, by inclusion/exclusion,
X
XG = f (∅) = (−1)|A|−|∅| g(A)
A⊇∅
X Y X |C|
= (−1)|A| xi
A⊆E(G) components C of λ(A) i≥1
X
|A|
= (−1) pλ(A) .
A⊆E(G)
(e) If you can solve this problem, please let me know so I can get some sleep. My former undergraduate
student Keeler Russell checked computationally in 2013 that there are no such pairs T, U with 24 or
fewer vertices, and more recently two undergraduates at Washington University (St. Louis) extended
this to 29. Matthew Morin, Jennifer Wagner and I proved that one can recover significant data about a
tree from its CSF, including its degree sequence and the number of vertices at every possible distance,
but this data does not determine the isomorphism type for trees with 11 or more vertices. Rosa Orel-
lana and Geoffrey Scott constructed an infinite family of pairs of unicyclic graphs (i.e., graphs that can
be made into trees by deleting one edge) with equal CSFs.
263
Chapter 10
For many combinatorial structures, there is a natural way of taking apart one object into two, or combining
two objects into one.
• Let G = (V, E) be a (simple, undirected) graph. For any W ⊆ V , we can break G into the two pieces
G|W and G|V \W . On the other hand, given two graphs, we can form their disjoint union G ∪· H.
• Let M be a matroid on ground set E. For any A ⊆ E, we can break M into the restriction M |A
(equivalently, the deletion of E \ A) and the contraction M/A. Two matroids can be combined into
one by taking the direct sum.
• Let P be a ranked poset. For any x ∈ P , we can extract the intervals [0̂, x] and [x, 1̂]. (Of course, we
don’t get every element of the poset this way.) Meanwhile, two graded posets P, Q can be combined
into one poset in many ways, such as Cartesian product (see Definition 1.1.12).
• Let α = (α1 , . . . , αℓ ) |= n. For 0 ≤ k ≤ ℓ, we can break α up into two sub-compositions α(k) =
(α1 , . . . , αk ), α(k) = (αk+1 , . . . , αℓ ). Of course, two compositions can be combined by concatenating
them.
In all these operations, there are lots of ways to split, but only one way to combine. Moreover, all the
operations are graded with respect to natural size functions on the objects: for instance, matroid direct sum
is additive on size of ground set and on rank.
Splitting Combining
|V (G|W )| + |V (G|V \W )| = |V (G)| |V (G ∪· H)| = |V (G)| + |V (H)|
|E(M |A )| + |E(M/A)| = |E(M ) |E(M ⊕ M ′ )| = |E(M )| + E(M ′ )|
r([0̂, x]) + r([x, 1̂]) = r(P ) r(P ⊕ Q) = r(P ) + r(Q)
(k)
|α(k) | + |α | = |α| r(αβ) = r(α) + r(β)
A Hopf algebra is a vector space H (over C, say) with two additional operations, a product µ : H ⊗ H → H
(which represents combining) and a coproduct ∆ : H → H⊗H which represents splitting. These operations
264
are respectively associative and coassociative, and they are compatible in a certain way. Technically, all this
data defines the slightly weaker structure of a bialgebra; a Hopf algebra is a bialgebra with an additional
map S : H → H, called the antipode. Most bialgebras that arise in combinatorics have a unique antipode
and thus a unique Hopf structure.
What is a C-algebra? It is a C-vector space A equipped with a ring structure. Its multiplication can be
thought of as a C-bilinear map
µ:A⊗A→A
that is associative, i.e., µ(µ(a, b), c) = µ(a, µ(b, c)). Associativity can be expressed as the commutativity of
the diagram
µ⊗Id
A⊗A⊗A A⊗A a⊗b⊗c ab ⊗ c
Id ⊗µ µ (10.1)
µ
A⊗A A a ⊗ bc abc
where I denotes the identity map. (Diagrams like this rely on the reader to interpret notation such as µ ⊗ I
as the only thing it could be possibly be; in this case, “apply µ to the first two tensor factors and tensor what
you get with [I applied to] the third tensor factor”.)
What then is a C-coalgebra? It is a C-vector space Z equipped with a C-linear comultiplication map
∆:Z →Z ⊗Z
that is coassociative, a condition defined by reversing the arrows in the previous diagram:
∆⊗Id
Z ⊗Z ⊗Z Z ⊗Z
Id ⊗∆ ∆ (10.2)
∆
Z ⊗Z Z
Just as an algebra has a unit, a coalgebra has a counit. To say what this is, let us diagramify the defining
property of the multiplicative unit 1A in an algebra A: it is the image of 1C under a map u : C → A such
that the diagram on the left commutes (where the top diagonal maps take a ∈ A to 1 ⊗ a or a ⊗ 1). Thus
a counit of a coalgebra is a map ε : Z → C such that the diagram on the right commutes (where the top
diagonal maps are projections).
A Z
u⊗Id Id ⊗u ε⊗Id Id ⊗ε
A⊗A Z ⊗Z
(10.3)
A bialgebra is a vector space B that has both a multiplication and a comultiplication, and such that multi-
plication is a coalgebra morphism and comultiplication is an algebra morphism. Both of these conditions
265
are expressible by commutativity of the diagram
∆⊗∆
B⊗B B⊗B⊗B⊗B
µ µ13 ⊗µ24 (10.4)
∆
B B⊗B
where µ13 ⊗ µ24 means the map that sends a ⊗ b ⊗ c ⊗ d to ac ⊗ bd (the subscripts refer to the positions of
the tensor factors).
Comultiplication takes some getting used to. As explained above, in combinatorial settings, one should
generally think of multiplication as putting two objects together, and comultiplication as taking an object
apart into two subobjects. A unit is a trivial object (putting it together with another object has no effect),
and the counit is the linear functional that picks off the coefficient of the unit.
Example 10.1.1 (The polynomial Hopf algebra). A simple example of a Hopf algebra is the polynomial ring
C[x]. It is an algebra in the usual way, and can be made into a coalgebra by the counit ε(f (x)) = f (0)
(equivalently, mapping every polynomial to its constant term) and the coproduct ∆(x) = 1 ⊗ x + x ⊗ 1.
Checking the bialgebra axioms is left as an exercise. ◀
Example 10.1.2 (The graph Hopf algebra). For n ≥ 0, let Gn be the set of formal C-linear combinations of
unlabeled simple graphs on n vertices (or if you prefer,
L of isomorphism classes [G] of simple graphs G,
but it is easier to drop the brackets), and let G = n≥0 n . Thus G is a graded vector space, which we
G
make into a C-algebra by defining µ(G ⊗ H) = G ∪· H, where ∪· denotes union under the assumption
V (G) ∩ V (H) = ∅.The unit is the unique graph K0 with no vertices (or, technically, the map u : C → G0
sending c ∈ C to cK0 ). Comultiplication in G is defined by
X
∆(G) = G|A ⊗ G|B .
·B
A,B: V (G)=A∪
As an illustration of how the compatibility condition (10.4) works, we will check it for G. To avoid “overfull
hbox” errors, set µ̃ = µ13 ⊗ µ24 . Then
X X
µ̃(∆ ⊗ ∆(G1 ⊗ G2 )) = µ̃ G1 |A1 ⊗ G1 |B1 ⊗ G2 |A2 ⊗ G2 |B2
A1 ∪
· B1 =V (G1 ) A2 ∪
· B2 =V (G2 )
X
= µ̃
G1 |A1 ⊗ G1 |B1 ⊗ G2 |A2 ⊗ G2 |B2
A1 ∪
· B1 =V (G1 )
A2 ∪
· B2 =V (G2 )
X
= (G1 |A1 ∪· G2 |A2 ) ⊗ (G1 |B1 ∪· G2 |B2 )
A1 ∪
· B1 =V (G1 )
A2 ∪
· B2 =V (G2 )
X
= (G1 ∪· G2 )|A ⊗ (G1 ∪· G2 )|B
· B=V (G1 ∪
A∪ · G2 )
= ∆(µ(G1 ⊗ G2 )).
Comultiplication in G is in fact cocommutative1 . Let sw be the “switching map” that sends a⊗b to b⊗a; then
commutativity and cocommutativity of multiplication and comultiplication on a bialgebra B are expressed
1 There are those who call this “mmutative”.
266
by the diagrams
sw sw
B⊗B B⊗B B⊗B B⊗B
µ µ ∆ ∆
B B
So cocommutativity means that ∆(G) is symmetric under switching; for the graph algebra this is clear
because A and B are interchangeable in the definition. ◀
Example 10.1.3 (Rota’s Hopf algebra of posets). For n ≥ 0, let Pn be the vector space of formal C-linear
combinations of isomorphism classes [P ] of finite graded posets P of rank n. Thus P0 and P1 are Lone-
dimensional (generated by the chains of lengths 0 and 1), but dim Pn = ∞ for n ≥ 2. We make P = n Pn
into a graded C-algebra by defining µ([P ] ⊗ [Q]) = [P × Q], where × denotes Cartesian product; thus
u(1) = •. Comultiplication is defined by
X
∆[P ] = [0̂, x] ⊗ [x, 1̂].
x∈P
Coassociativity is checked by the following calculation, which should remind you of the proof of associa-
tivity of convolution in the incidence algebra of a poset (Prop. 2.1.2):
!
X
∆ ⊗ I(∆(P )) = ∆ ⊗ I [0̂, x] ⊗ [x, 1̂]
x∈P
X
= ∆([0̂, x]) ⊗ [x, 1̂]
x∈P
X X
= [0̂, y] ⊗ [y, x] ⊗ [x, 1̂]
x∈P y∈[0̂,x]
X
= [0̂, y] ⊗ [y, x] ⊗ [x, 1̂]
x≤y∈P
X X
= [0̂, y] ⊗ [y, x] ⊗ [x, 1̂]
y∈P x∈[y,1̂]
X
= [0̂, y] ⊗ ∆([y, 1̂]) = I ⊗ ∆(∆(P )).
y∈P
This Hopf algebra is commutative, but not cocommutative; the switching map does not fix ∆(P ) unless P
is self-dual. ◀
Example 10.1.4 (The Hopf algebra of matroids). For n ≥ 0, let Mn be the vector space of formal C-linear
combinations of isomorphism classes [M L] of finite matroids M on n elements. Here dim P0 = 1 and
dim Pn < ∞ for every n. We make M = n Mn into a graded C-algebra by defining µ([P ]⊗[Q]) = [P ⊗Q].
The trivial matroid (with empty ground set) is the multiplicative identity. Note that multiplication is com-
mutative. Letting E denote the ground set of M , we define comultiplication by
X
∆[M ] = M |A ⊗ M/A.
A⊆E
Coassociativity is essentially a consequence of the compatibility of deletion and contraction (Prop. 3.8.2).
Note that the coproduct is not cocommutative. ◀
267
This is a good place to introduce what is known as Sweedler notation. Often, it is highly awkward to notate
all the summands in a coproduct, particularly if we are trying to prove general facts about Hopf algebra.
The Sweedler notation for a coproduct is
X
∆(h) = h1 ⊗ h2
which should be read as “the coproduct of h is a sum of a bunch of tensors, each of which has a first element
and a second element.” This notation looks dreadfully abusive at first, but in fact it is incredibly convenient,
is unambiguous if used properly, and one soon discovers that any other way of doing things would be
worse (imagine having to conjure an index set out of thin air and deal with a lot of double subscripts just to
write down a coproduct). Sweedler notation iterates well; for example, we could write
X
∆2 (h) = (Id ⊗∆)(∆(h)) = (∆ ⊗ Id)(∆(h)) = h1 ⊗ h2 ⊗ h3
For example, clearly ∆(c) = c = c ⊗ 1 = 1 ⊗ c for any scalar c. Moreover, for every k, we have
k
X k
X
hk (x, y) = hj (x)hk−j (y), ek (x, y) = ej (x)ek−j (y)
j=0 j=0
and therefore
k
X k
X
∆(hk ) = hj ⊗ hk−j , ∆(ek ) = ej ⊗ ek−j .
j=0 j=0
◀
Definition 10.1.6. A Hopf algebra is a bialgebra H with a antipode S : H → H, which satisfies the commu-
tative diagram
S⊗Id
H⊗H H⊗H
∆ µ
H H (10.5)
ε u
P
In other words, to calculate the antipode of something, comultiplyP it to get ∆g = g1 ⊗ g2 . Now hit every
first tensor factor with S and then multiply it out again to obtain S(g1 ) · g2 . If you started with the unit
268
then this should be 1, while if you started with any other homogeneous object then you get 0. This enables
calculating the antipode recursively. For example, in QSym:
Lemma 10.1.7 (Humpert, Prop 1.4.4). Let B be a bialgebra that is graded and connected, i.e., the 0th graded piece
has dimension 1 as a vector space. Let n > 0 and let h ∈ Hn . Then
X
∆(h) = h ⊗ 1 + h1 ⊗ h2 + 1 ⊗ h
where the Sweedler-notation sum contains only elements of degrees strictly between 0 and n.
Proof. Refer to the diagrams for the unit and counit (10.3). In particular, the right-hand triangle gives
h1 ⊗ ε(h2 ) = h. So certainly one of the summands must have h1 ∈ Hn , but then h2 ∈ H0 . Since H0 ∼
=C
P
we may as well group all those summands together; they must sum to h ⊗ 1. Meanwhile, the left-hand
triangle says that grouping together all the summands of bidegree 0, n gives 1 ⊗ h.
Proposition 10.1.8. Let B be a connected and graded bialgebra. Then the commutative diagram (10.5) defines a
unique antipode S : B → B, and thus B can be made into a Hopf algebra in a unique way.
Combinatorics features lots of graded connected bialgebras (such as all those we have seen so far), so this
proposition gives us a Hopf algebra structure “for free”.
There is a general recipe for the antipode, known as Takeuchi’s formula [Tak71]. Let π : H → H be the map
that kills H0 and fixes each positive graded piece pointwise. Then
X
S = uε + (−1)k µk−1 π ⊗k ∆k−1 , (10.6)
k≥1
i.e., X X
S(h) = u(ε(h)) − π(h) + π(h1 )π(h2 ) − π(h1 )π(h2 )π(h3 ) + · · ·
However, there is a lot of cancellation in this sum, making it impractical for looking at specific Hopf alge-
bras. Therefore, one of the first things one wants in studying a particular Hopf algebra is to find a cleaner
formula for the antipode. An excellent example is the Hopf algebra of symmetric functions, in which the
antipode is the involution ω interchanging eλ and hλ (Proposition 9.5.3). The proof is left as an exercise
(Problem 10.3).
269
10.2 Characters
A character on a Hopf algebra H is a C-linear map ζ : H → C that is multiplicative, i.e., ζ(1H ) = 1C and
ζ(h · h′ ) = ζ(h)ζ(h′ ). For example, if H is the graph Hopf algebra, then we can define a character by
(
1 if G has no edges,
ζ(G) = (10.7)
0 if G has one or more edges,
for a graph G, and then extending by linearity to all of G. This map is multiplicative (because G · H has an
edge iff either G or H does); it also looks kind of like a silly map. However, the reason this is interesting
is that characters can be multiplied
P together. The multiplication is called convolution product, defined as
follows: if h ∈ H and ∆(h) = h1 ⊗ h2 in Sweedler notation, then
X
(ζ ∗ η)(h) = ζ(h1 )η(h2 ).
One can check that convolution is associative; the calculation resembles checking that the incidence algebra
of a poset is an algebra. The counit ε is a two-sided identity for convolution, i.e., ζ ∗ ε = ε ∗ ζ = ζ for all
characters ζ. Moreover, the definition (10.5) of the antipode implies that
ζ ∗ (ζ ◦ S) = ε
(check this too). Therefore, the set of all characters forms a group.
Why would you want to convolve characters? Consider the graph Hopf algebra with the character ζ, and
let k ∈ N. The kth convolution power of ζ is given by
X
ζ k (G) = ζ ∗ · · · ∗ ζ (G) = ζ(G|V1 ) · · · ζ(G|Vk )
| {z }
k times V (G)=V1 ∪
· ···∪
· Vk
(
X 1 if V1 , . . . , Vk are all cocliques,
=
V (G)=V1 ∪
· ···∪
· Vk
0 otherwise.
(recall that a coclique is a set of vertices of which no two are adjacent). In other words, ζ n (G) counts the
number of functions f : V → [k] so that f (x) ̸= f (y) whenever x, y are adjacent. But such a thing is precisely
a proper k-coloring! I.e.,
ζ n (G) = p(G; k)
where p is the chromatic polynomial (see Section 4.4). This turns out to be true as a polynomial identity
in k — for instance, ζ −1 (G) is the number of acyclic orientations. One can even view the Tutte polynomial
k
T (G; x, y) as a character τx,y (G) with parameters x, y; it turns out that τx,y (G) is itself a Tutte polynomial
evaluation — see Brandon Humpert’s Ph.D. thesis [Hum11].
A combinatorial Hopf algebra, or CHA, is a pair (H, ζ), where H is a graded connected Hopf algebra and ζ
Φ
→ (H′ , ζ ′ ) that is an algebra and coalgebra morphism
is a character. A morphism of CHA’s is a map (H, ζ) −
′ ′
and satisfies ζ ◦ Φ = Φ ◦ ζ .
Example 10.2.1. The binomial Hopf algebra is the ring of polynomials C[x], equipped with the coproduct
generated by ∆(x) = x ⊗ 1 + 1 ⊗ x. To justify the name, note that
n
X n
∆(xn ) = ∆(x)n = (x ⊗ 1 + 1 ⊗ x)n = xk ⊗ xn−k .
k
k=0
270
This is extended linearly, so that ∆(f (x)) = f (∆(x)) for any polynomial f . The counit is ε(f ) = f (0), and
the antipode is given by S(xk ) = (−1)k xk (check this). We make it into a CHA by endowing it with the
character ε1 (f ) = f (1).
Pζ : (H, ζ) → (C[x], ε1 )
For example, if H is the graph algebra and ζ the characteristic function of edgeless graphs (10.7), then Pζ is
the chromatic polynomial. ◀
Example 10.2.2. The ring QSym of quasisymmetric functions can be made into a Hopf algebra as follows.
Let α = (α1 , . . . , αk ) be a composition; then
k
X
∆Mα = M(α1 ,...,αj ) ⊗ M(αj+1 ,...,αk ) .
j=0
One can check (Problem 10.3) that the Hopf algebra of symmetric functions described in Example 10.1.5 is
a Hopf subalgebra of QSym; that is, this coproduct restricts to the one defined earlier on Λ. We then endow
QSym with the character ζQ defined on the level of power series by ζQ (x1 ) = 1 and ζQ (xj ) = 0 for j ≥ 2;
equivalently, (
1 if α has at most one part,
ζQ (Mα ) =
0 otherwise.
One of the main theorems about CHAs, due to Aguiar, Bergeron and Sottile [ABS06], is that (QSym, ζQ ) is a
terminal object in the category of CHAs, i.e., every CHA (H, ζ) admits a canonical morphism to (QSym, ζ).
For the graph algebra, this morphism is the chromatic symmetric function; for the matroid algebra, it is the
Billera-Jia-Reiner invariant. ◀
Hopf monoids are a more recent area of research. One exhaustive reference is the book by Aguiar and
Mahajan [AM10]; more accessible introductions (and the main sources for these notes) include Klivans’
talk slides [Kli] and the preprint by Aguiar and Ardila [AA17]. One of the ideas behind Hopf monoids is to
work with labeled rather than unlabeled objects.
First, we need a set H[I] for every finite set I. One should think of H[I] as the vector space spanned by com-
binatorial objects of a certain ilk, with I as the labeling set. (For example, graphs with vertices I, matroids
with ground set I, linear orderings of I, polyhedra in RI , etc.) Every bijection π : I → I ′ should induce a
linear isomorphism H[π] : H[I] → H[I ′ ], which should be thought of as relabeling, and the association of
H[π] with π is functorial2 . A functor H with these properties is called a vector species. Moreover, we require
that dim H[∅] = 1, and we identify a particular nonzero element of H[∅] as the “trivial object”.
2 This is a fancy way of saying that it obeys some completely natural identities: H[Id ] = Id
I H[I] and H[π ◦ σ] = H[π] ◦ H[σ]. Don’t
worry too much about it.
271
Then, we need to have multiplication and comultiplication maps for every decomposition I = A ∪· B:
µA,B ∆A,B
H[A] ⊗ H[B] −−−→ H[I] and H[I] −−−→ H[A] ⊗ H[B]. (10.8)
These are subject to a whole lot of conditions. The most important of these are labeled versions of associa-
tivity, coassociativity, and compatibility:
µI,J ⊗IdK
H[I] ⊗ H[J] ⊗ H[K] H[I] ⊗ H[J ∪· K]
∆I,J ⊗IdK
H[I] ⊗ H[J] ⊗ H[K] H[I] ⊗ H[J ∪· K]
∆I,J ⊗∆K,L
H[I ∪· J] ⊗ H[K ∪· L] H[I] ⊗ H[J] ⊗ H[K] ⊗ H[L]
µI∪
· J,K∪
·L (µI,K ⊗µJ,L )◦τ (compatibility), (10.11)
∆I∪
· K,J∪
·L
H[I ∪· J ∪· K ∪· L] H[I ∪· K] ⊗ H[J ∪· L]
Note that instead of defining a single coproduct as the sum over all possible decompositions A, B (as in the
Hopf algebra setup), we are keeping the different decompositions separate.
In many cases, the operations can be defined on the level of individual combinatorial objects. In other
words, we start with a set species h — a collection of sets h[I] indexed by finite sets I, subject to the
conditions that any bijection I → I ′ naturally induces a bijection h[I] → h[I ′ ], define multiplication and
comultiplication operations
µA,B ∆A,B
h[A] × h[B] −−−→ h[I] and h[I] −−−→ h[A] × h[B]
(in contrast to eqrefvector-species-product-coproduct, these are Cartesian products of sets rather than ten-
sor products of vector spaces), then define a vector species H by setting H[I] = kh[I], and define multi-
plication and comultiplication on H by linear extension. Such a Hopf monoid is called linearized. This is
certainly a very natural kind of Hopf monoid, but not all the Hopf monoids we care about come from a set
species in this way.
Example 10.3.1. Let ℓ [I] denote the set of linear orders on a finite set I, which we can think of as bijections
w : [n] → I (and represent by the sequence w(1), . . . , w(n)). Given a decomposition I = A ∪· B, the most
obvious way to define product and coproduct on the set species ℓ is by concatenation and restriction. For
instance, if A = {a, b, c} and B = {p, q, r, s}, then
Linearizing this setup produces the Hopf monoid of linear orders L = kℓℓ. ◀
272
Example 10.3.2. Let m[I] denote the set of matroids with ground set I, with product and coproduct defined
setwise by
µ(M1 , M2 ) = M1 ⊕ M2 , ∆A,B (M ) = (M |A , M/A).
The linearized Hopf monoid M = km is a labeled analogue of the matroid Hopf algebra M described in
Example 10.1.4. ◀
Multiplication and comultiplication can be iterated. For any set composition A (i.e., an ordered list A =
A1 | . . . |An whose disjoint union is I), there are maps
n n
O µA ∆A
O
H[Ai ] −−→ H[I] and H[I] −−→ H[Ai ]
i=1 i=1
that are well defined by associativity and coassociativity. (For set species, replace tensor products with
Cartesian products.) For example, if A = (I, J, K) then we can define µA by either traveling south then
east, or east then south, in (10.9) — we get the same answer in both cases.
The antipode in a Hopf monoid H is the following collection of maps SI : H[I] → H[I] given by the Takeuchi
formula: for x ∈ H[I], (
x if I = ∅,
S(x) = SI (x) = P n
(10.12)
A|=I (−1) µA (∆A (x)) if I ̸= ∅.
Here A |= I means that A runs over all set compositions of I with nonempty parts (in particular, there are only
finitely many summands). As in the Hopf algebra setting, this formula typically has massive cancellation,
so in order to study a particular Hopf monoid it is desirable to find a cancellation-free formula.
Example 10.3.3. Let us calculate some antipodes in L. The trivial ordering on ∅ is trivially fixed by S, while
for a singleton set I = {a} we have S(a) = −a (the Takeuchi formula has only one term, corresponding to
the set partition of I with one block). For ab ∈ L[{a, b}] we have
S(ab) = −µ12 (∆12 (ab)) + µ1|2 (∆1|2 (ab)) + µ2|1 (∆2|1 (ab))
= −ab + (a)(b) + (b)(a)
= −ab + ab + ba = ba,
while for abc ∈ L[I] the antipode is calculated by the following table:
It is starting to look suspiciously as though SI = (−1)I rev, where rev denotes the map that reverses order-
ing. In fact this is the case (proof left as an exercise). ◀
273
Material to be written: duality, L∗ , generalized permutahedra and the Aguilar-Ardila antipode calculation,
...
10.4 Exercises
Problem 10.1. Show that in a Hopf algebra one has ∆ ◦ µ = Id. (This is known as “Enrique’s lemma” at
KU.) An immediate corollary is that product and coproduct are injective and surjective, respectively.
Solution: To be written. Mark writes: “This is Corollary 8.38 (pg 264) in ”Monoidal Functors, Species,
and Hopf Algebras” The top left diagram. It also shows up as Proposition 7 (pg 26) in ”Hopf Monoids in
Species”. I think the proof of this using the compatibility axiom is a reasonable exercise.
Problem 10.2. Confirm that the polynomial Hopf algebra (Example 10.1.1) satisfies (10.2) and (10.4), and
determine its antipode.
Solution: To be written
Problem 10.3. Confirm that the symmetric functions Λ form a Hopf subalgebra of the quasi-symmetric
functions QSym, as asserted in Example 10.2.2. Then show that the antipode on Λ is the involution ω
interchanging eλ and hλ (Proposition 9.5.3).
Solution: To be written. For the antipode, our calculation of ∆(hk ) says that
k
(
X 1 if k = 0,
S(hj )hk−j =
j=0
0 if k > 0
and comparing with the Jacobi-Trudi relations (see §9.5) we see that S(hk ) = (−1)k ek , i.e., S = (−1)k ω.
Problem 10.4. Let E(M ) denote the ground set of a matroid M , and call |E(M )| the order of M . Let Mn be
L space of formal C-linear combinations of isomorphism classes [M ] of matroids M of order n. Let
the vector
M = n≥0 Mn . Define a graded multiplication on M by [M ][M ′ ] = [M ⊕ M ′ ] and a graded comultiplica-
tion by X
∆[M ] = [M |A ] ⊗ [M/A]
A⊆E(M )
where M |A and M/A denote restriction and contraction respectively. Check that these maps make M into
a graded bialgebra, and therefore into a Hopf algebra by Proposition 10.1.8.
Solution: To be written
Problem 10.5. Prove that the Billera–Jia–Reiner invariant defines a Hopf algebra morphism M → QSym.
Solution: To be written
Problem 10.6. Prove that the antipode in L is indeed given by SI = (−1)I rev, as in Example 10.3.3.
Solution: To be written
274
Chapter 11
More Topics
The main theorem of this section is the Max-Flow/Min-Cut Theorem of Ford and Fulkerson. Strictly speak-
ing, it probably belongs to graph theory or combinatorial optimization rather than algebraic combinatorics,
but it is a wonderful theorem and has applications to posets and algebraic graph theory, so I can’t resist
including it.
Definition 11.1.1. A network N consists of a directed graph (V, E), two distinct vertices s, t ∈ V (called the
source and sink respectively), and a capacity function c : E → R≥0 .
Throughout this section, we will fix the symbols V , E, s, t, and c for these purposes. We will assume that
the network has no edges into the source or out of the sink.
A network is supposed to model the flow of “stuff”—data, traffic, liquid, electrical current, etc.—from s to t.
The capacity of an edge is the maximum amount of stuff that can flow through it (or perhaps the amount of
stuff per unit time). This is a general model that can be specialized to describe cuts, connectivity, matchings
and other things in directed and undirected graphs. This interpretation is why we exclude edges into s or
out of t; we will see later why this assumption is in fact justified.
If c(e) ∈ N for all e ∈ E, we say the network is integral. In what follows, we will only consider integral
networks.
a 1 b
1 1
s 2
t
1 1
c 1
d
275
Definition 11.1.2. A flow on N is a function f : E → N that satisfies the capacity constraints
where X X
f − (v) = f (e), f + (v) = f (e).
e=−
→
uv e=−→
vw
The number f (e) represents the amount of stuff flowing through e. That amount is bounded by the capacity
of that edge, hence the constraints (11.1). Meanwhile, the conservation constraints say that stuff cannot
accumulate at any internal vertex of the network, nor can it appear out of nowhere.
The max-flow problem is to find a flow of maximum value. The dual problem is the min-cut problem,
which we now describe.
Definition 11.1.3. Let N be a network. Let S, T ⊆ V with S ∪ T = V , S ∩ T = ∅, s ∈ S, and t ∈ T . The
corresponding cut is
[S, T ] = {−→ ∈ E : x ∈ S, y ∈ T }
xy
and the capacity of that cut is X
c(S, T ) = c(e).
e∈[S,T ]
A cut can be thought of as a bottleneck through which all stuff must pass. For example, in the network of
→
− − → → −
Figure 11.1, we could take S = {s, a, c}, T = {b, d, t}, so that [S, T ] = {ab, ad, cd}, and c(S, T ) = 1+2+1 = 4.
The min-cut problem is to find a cut of minimum capacity. This problem is certainly feasible, since there
are only finitely many cuts and each one has finite capacity.
X X
For A ⊆ V , define f − (A) = f (e), f + (A) = f (e).
e∈[Ā,A] e∈[A,Ā]
The proof (which requires little more than careful bookkeeping) is left as an exercise.
276
The inequality (11.3c) is known as weak duality; it says that the maximum value of a flow is less than or
equal to the minimum capacity of a cut. (Strong duality would say that equality holds.)
Suppose that there is a path P from s to t in which no edge is being used to its full capacity. Then we can
increase the flow along every edge on that path, and thereby increase the value of the flow by the same
amount. As a simple example, we could start with the zero flow f0 on the network of Figure 11.1 and
increase flow by 1 on each edge of the path sadt; see Figure 11.2.
a 10 b a 10 b
10 10 11 10
s 20 s 21
t t
10 10 10 11
c 10 c 10
d d
|f0 | = 0 |f1 | = 1
The problem is that there can exist flows that cannot be increased in this elementary way — but nonetheless
are not maximum. The flow f1 of Figure 11.2 is an example. In every path from s to t, there is some edge e
with f (e) = c(e). However, it easy to construct a flow of value 2:
a 11 b
11 11
s 20
t
11 11
c 11
d
|f2 | = 2
Figure 11.3: A better flow that cannot be obtained from f1 in the obvious way.
Fortunately, there is a more general way to increase the value of a flow. The key idea is that flow along an
edge −→ can be regarded as negative flow from y to x. Accordingly, all we need is a path from s to t in which
xy
each edge e is either pointed forward and has f (e) < c(e), or is pointed backward and has f (e) > 0. Then,
increasing flow on the forward edges and decreasing flow on the backward edges will increase the value of
the flow. This is called an augmenting path for f .
The Ford-Fulkerson Algorithm is a systematic way to construct a maximum flow by looking for augment-
ing paths. The wonderful feature of the algorithm is that if a flow f has no augmenting path, the algorithm
will automatically find a cut of capacity equal to |f | — thus certifying immediately that the flow is maxi-
mum and that the cut is minimum.
277
a 10 b a 11 b
11 10 11 11
s 21 s 20
t t
10 11 11 11
c 10 c 11
d d
|f1 | = 1 |f2 | = 2
Figure 11.4: Exploiting the augmenting path scdabt for f1 . The flow is increased by 1 on each of the “for-
ward” edges sc, cd, ab, bt and decreased by 1 on the “backward” edge da to obtain the improved flow f2 .
P : x0 = s, e1 , x1 , e2 , x2 , . . . , xn−1 , en , xn = t
By integrality and induction, all tolerances are integers and all flows are integer-valued. In particular, each
iteration of the loop increases the value of the best known flow by 1. Since the value of every flow is
bounded by the minimum capacity of a cut (by weak duality), the algorithm is guaranteed to terminate in
a finite number of steps. (By the way, Step 1 of the algorithm can be accomplished efficiently by a slight
modification of, say, breadth-first search.)
The next step is to prove that this algorithm actually works. That is, when it terminates, it will have com-
puted a flow of maximum possible value.
Proposition 11.1.5. Suppose that f is a flow that has no augmenting path. Let
Then s ∈ S, t ∈ T , and c(S, T ) = |f |. In particular, f is a maximum flow and [S, T ] is a minimum cut.
Proof. Note that t ̸∈ S precisely because f has no augmenting path. Applying (11.3b) gives
X X X
|f | = f + (S) − f − (S) = f (e) − f (e) = f (e).
e∈[S,S̄] e∈[S̄,S] e∈[S,S̄]
But f (e) = c(e) for every e ∈ [S, T ] (otherwise S would be bigger than what it actually is), so this last
quantity is just c(S, T ). The final assertion follows by weak duality.
We have proven:
278
Theorem 11.1.6 (Max-Flow/Min-Cut Theorem for Integral Networks (“MFMC”)). For every integral network
N , the maximum value of a flow equals the minimum value of a cut.
In light of this, we will call the optimum of both the max-flow and min-cut problems the value of N , written
|N |. In fact MFMC holds for non-integral networks as well, although the Ford-Fulkerson algorithm may
not work in that case (the flow value might converge to |N | without ever reaching it.)
Definition 11.1.7. Let N be a network. A flow f in N is acyclic if, for every directed cycle C in N (i.e., every
set of edges x1 → x2 → · · · → xn → x1 ), there is some e ∈ C for which f (e) = 0. The flow f is partitionable
if there is a collection of s, t-paths P1 , . . . , P|f | such that for every e ∈ E,
f (e) = #{i : e ∈ Pi }.
(Here “s, t-path” means “path from s to t”.) In this sense f can be regarded as the “sum” of the paths Pi ,
each one contributing a unit of flow.
Proposition 11.1.8. Let N be a network. Then:
1. For every flow in N , there exists an acyclic flow with the same value. In particular, N admits an acyclic flow
with |f | = |N |.
2. Every acyclic integral flow is partitionable.
Proof. Suppose that some directed cycle C has positive flow on every edge. Let k = min{f (e) : e ∈ C}.
Define f˜ : E → N by (
f (e) − k if e ∈ C,
f˜(e) =
f (e) if e ̸∈ C.
Given a nonzero acyclic flow f , find an s, t-path P1 along which all flow is positive. Decrement the flow
on each edge of P1 ; doing this will also decrement |f |. Now repeat this for an s, t-path P2 , etc. When the
resulting flow is zero, we will have partitioned f into a collection of s, t-paths of cardinality |f |.
Remark 11.1.9. This discussion justifies our earlier assumption that there are no edges into the source or
out of the sink, since every acyclic flow must be zero on all such edges. Therefore, deleting those edges
from a network does not change the value of its maximum flow.
This result has many applications in graph theory: Menger’s theorems, the König-Egerváry theorem, etc.
The basic result in this area is Dilworth’s Theorem, which resembles the Max-Flow/Min-Cut Theorem (and
can indeed be derived from it; see the exercises).
Definition 11.2.1. A chain cover of a poset P is a collection of chains whose union is P . The minimum size
of a chain cover is called the width of P .
279
Theorem 11.2.2 (Dilworth’s Theorem). Let P be a finite poset. Then
width(P ) = m(P ).
Proof. The “≥” direction is clear, because if A is an antichain, then no chain can meet A more than once, so
P cannot be covered by fewer than |A| chains.
For the more difficult “≤” direction, we induct on n = |P |. The result is trivial if n = 1 or n = 2.
Let Y be the set of all minimal elements of P , and let Z be the set of all maximal elements. Note that Y
and Z are both antichains. First, suppose that no set other than Y or Z is a maximum1 antichain; dualizing
if necessary, we may assume |Y | = m(P ). Let y ∈ Y and z ∈ Z with y ≤ z. Let P ′ = P \ {y, z}’; then
m(P ′ ) = |Y | − 1. By induction, width(P ′ ) ≤ |Y | − 1, and taking a chain cover of P ′ and tossing in the chain
{y, z} gives a chain cover of P of size |Y |.
Then
• P + , P − ̸= A (otherwise A equals Z or Y ).
• P + ∪ P − = P (otherwise A is contained in some larger antichain).
• P + ∩ P − = A (otherwise A isn’t an antichain).
So P + and P − are posets smaller than P , each of which contains A as a maximum antichain. By induction,
each P ± has a chain cover of size |A|. So for each a ∈ A, there is a chain Ca+ ⊆ P + and a chain Ca− ⊆ P −
with a ∈ Ca+ ∩ Ca− , and
Ca ∩ Ca− : a ∈ A}
+
If we switch “chain” and “antichain”, then Dilworth’s theorem remains true and becomes a much easier
result.
Proposition 11.2.3 (Mirsky’s Theorem). In any finite poset, the minimum size of an antichain cover equals the
maximum size of an chain.
Proof. For the ≥ direction, if C is a chain and A is an antichain cover, then no antichain in A can contain
more than one element of C, so |A| ≥ |C|. On the other hand, let
then {Ai } is an antichain cover whose cardinality equals the length of the longest chain in P .
There is a marvelous common generalization of Dilworth’s and Mirsky’s Theorems due to Curtis Greene
and Daniel Kleitman [GK76, Gre76]. An excellent source on this topic, including multiple proofs, is the
survey article [BF01] by Thomas Britz and Sergey Fomin.
1 I.e., a chain of size m(P ) — not merely a chain that is maximal with respect to inclusion, which might have smaller cardinality.
280
Theorem 11.2.4 (Greene-Kleitman). Let P be a finite poset. Define two sequences of positive integers
λ = (λ1 , λ2 , . . . , λℓ ), µ = (µ1 , µ2 , . . . , µm )
by
λ1 + · · · + λk = max |C1 ∪ · · · ∪ Ck | : Ci ⊆ P chains ,
µ1 + · · · + µk = max |A1 ∪ · · · ∪ Ak | : Ai ⊆ P disjoint antichains .
Then:
1. λ and µ are both partitions of |P |, i.e., weakly decreasing sequences whose sum is |P |.
2. λ and µ are conjugates (written µ = λ̃): the row lengths of λ are the column lengths in µ, and vice versa.
Note that Dilworth’s Theorem is just the special case µ1 = ℓ. As an example, the poset with Hasse diagram
has
λ = (3, 2, 2, 2) = and µ = (4, 4, 1) = = λ̃.
How many different necklaces can you make with four blue, two green, and one red bead?
It depends what “different” means. The second necklace can be obtained from the first by rotation, and the
third by reflection, but the fourth one is honestly different from the first two.
If we just wanted to count the number of ways to permute four blue, two green, and one red beads, the
answer would be the multinomial coefficient
7 7!
= = 105.
4, 2, 1 4! 2! 1!
However, what we are really trying to count is orbits under a group action.
281
Let G be a group and X a set. An action of G on X is a group homomorphism α : G → SX , the group of
permutations of X.
Equivalently, an action can also be regarded as a map G × X → X, sending (g, x) to gx, such that
• IdG x = x for every x ∈ X (where IdG denotes the identity element of G);
• g(hx) = (gh)x for every g, h ∈ G and x ∈ X.
To go back to the necklace problem, we now see that “same” really means “in the same orbit”. In this
case, X is the set of all 105 necklaces, and the group acting on them is the dihedral group D7 (the group of
symmetries of a regular heptagon). The number we are looking for is the number of orbits of D7 .
Lemma 11.3.1. Let x ∈ X. Then |Ox ||Sx | = |G|.
Proof. The element gx depends only on which coset of Sx contains g, so |Ox | is the number of cosets, which
is |G|/|Sx |.
Proposition 11.3.2. [Burnside’s Theorem] The number of orbits of the action of G on X equals the average number
of fixed points:
1 X
#{x ∈ X : gx = x}
|G|
g∈G
Proof. For a sentence P , let χ(P ) = 1 if P is true, or 0 if P is false (the “Garsia chi function”). Then
X 1 1 X
Number of orbits = = |Sx |
|Ox | |G|
x∈X x∈X
1 XX
= χ(gx = x)
|G|
x∈X g∈G
1 XX 1 X
= χ(gx = x) = #{x ∈ X : gx = x}.
|G| |G|
g∈G x∈X g∈G
282
Therefore, the number of orbits is
105 + 7 · 3 126
= = 9,
|D7 | 14
which is much more pleasant than trying to count them directly. ◀
Example 11.3.4. Suppose we wanted to find the number of orbits of 7-bead necklaces with 3 colors, without
specifying how many times each color is to be used.
k 7 + 7k 4 + 6k
. (11.4)
14
◀
As this example indicates, it is helpful to look at the cycle structure of the elements of G, or more precisely
on their images α(g) ∈ SX .
Proposition 11.3.5. Let X be a finite set, and let α : G → SX be a group action. Color the elements of X with k
colors, so that G also acts on the colorings.
1. For g ∈ G, the number of fixed points of the action of g is k ℓ (g), where ℓ(g) is the number of cycles in the
disjoint-cycle representation of α(g).
2. Therefore,
1 X ℓ(g)
#equivalence classes of colorings = k . (11.5)
|G|
g∈G
Let’s rephrase Example 11.3.4 in this notation. The identity has cycle-shape 1111111 (so ℓ = 7); each of the
six reflections has cycle-shape 2221 (so ℓ = 4); and each of the seven rotations has cycle-shape 7 (so ℓ = 1).
Thus (11.4) is an example of the general formula (11.5).
Example 11.3.6. How many ways are there to k-color the vertices of a tetrahedron, up to moving the tetra-
hedron around in space?
Here X is the set of four vertices, and the group G acting on X is the alternating group on four elements.
This is the subgroup of S4 that contains the identity, of cycle-shape 1111; the eight permutations of cycle-
shape 31; and the three permutations of cycle-shape 22. Therefore, the number of colorings is
k 4 + 11k 2
.
12
◀
283
11.4 Grassmannians
A standard reference for everything in this and the following section is Fulton [Ful97].
One motivation for the combinatorics of partitions and tableaux comes from classical enumerative geomet-
ric questions like this:
The Four-Lines Problem: Let there be given four lines L1 , L2 , L3 , L4 in R3 in general position. How many
lines M meet each of L1 , L2 , L3 , L4 nontrivially?
To a combinatorialist, “general position” means “all pairs of lines are skew, and their direction vectors are
as linearly independent as possible — that is, the matroid they represent is U3 (4).” To a probabilist, it means
“choose the lines randomly according to some reasonable measure on the space of all lines.” So, what does
the space of all lines look like?
In general, if V is a vector space over a field k (which we will henceforth take to be R or C), and 0 ≤
k ≤ dim V , then the space of all k-dimensional vector subspaces of V is called the Grassmannian Gr(k, V ).
(Warning: this notation varies considerably from source to source.) As we will see, Gr(k, V ) has many nice
properties:
where ≤ is the usual partial order on Young’s lattice (i.e., containment of Ferrers diagrams).
• When k = C, the Poincaré polynomial of Gr(k, Cn ), i.e., the Hilbert series of the cohomology ring of
n C ), is the rank-generating function for the graded poset Yk,n , namely, the q-binomial coefficient
n 2
Gr(k,
k q (see Problem 2.7(c)).
To accomplish all this, we need some way to describe points of the Grassmannian. For as long as possible,
we won’t worry about the ground field.
Let W ∈ Gr(k, kn ); that is, W is a k-dimensional subspace of V = kn . We can describe W as the column
2 If these terms don’t make sense, here is a sketch of what you need to know. The cohomology ring H ∗ (X) = H ∗ (X; Q) of a space
X is just some ring that is a topological invariant of X. If X is a reasonably civilized space — say, a compact finite-dimensional real or
complex manifold, or a finite simplicial complex — then H ∗ (X) is a graded ring H 0 (X) ⊕ H 1 (X) ⊕ · · · ⊕ H d (X), where d = dim X,
and each graded piece H i (X) is a finite-dimensional Q-vector space. The Poincaré polynomial records the dimensions of these vector
spaces as a generating function:
Xd
Poin(X, q) = dimQ H i (X) q i .
i=0
For lots of spaces, this polynomial has a nice combinatorial formula. For instance, take X = RP d (real projective d-space). It turns out
that H ∗ (X) ∼= Q[z]/(z n+1 ). Each graded piece H i (X), for 0 ≤ i ≤ d, is a 1-dimensional Q-vector space (generated by the monomial
xi ), and Poin(X, q) = 1 + q + q 2 + · · · + q d = (1 − q d+1 )/(1 − q). In general, if X is a compact orientable manifold, then Poincaré
duality implies (among other things) that Poin(X, q) is a palindrome.
284
space of a n × k matrix M of full rank:
m11 ··· m1k
.. .. .
M = . .
mn1 ··· mnk
However, the Grassmannian is not simply the space Z of all such matrices, because many different matrices
can have the same column space. Specifically, any invertible column operation on M leaves its column
space unchanged. On the other hand, every matrix whose column space is W can be obtained from M by
some sequence of invertible column operations; that is, by multiplying on the right by some invertible k × k
matrix. Accordingly, it makes sense to write
That is, the k-dimensional subspaces of kn can be identified with the orbits of Z under the action of the
general linear group GLk (k). In fact, as one should expect from (11.6),
where “dim” means dimension as a manifold over k; note that dim Z = nk because Z is a dense open subset
of kn×k . (Technically, this dimension calculation does not follow from (11.6) alone; you need to know that
the action of GLk (k) on Z is suitably well-behaved. Nevertheless, we will soon be able to calculate the
dimension of Gr(k, kn ) more directly.)
We now want to find a canonical representative for each GLk (k)-orbit. In other words, given W ∈ Gr(k, kn ),
we want the “nicest” matrix whose column space is W . How about the reduced column-echelon form?
Basic linear algebra says that we can pick any matrix with column space W and perform Gauss-Jordan
elimination on its columns, ending up with a uniquely determined matrix M = M (W ) with the following
properties:
• colspace M = W .
• The top nonzero entry of each column of M (the pivot in that column) is 1.
• Let pi be the row in which the ith column has its pivot. Then 1 ≤ p1 < p2 < · · · < pk ≤ n.
• Every entry below a pivot of M is 0, as is every entry to the right of a pivot.
• The remaining entries of M (i.e., other than the pivots and the 0s just described) can be anything
whatsoever, depending on what W was in the first place.
For example, if n = 4 and k = 2, then M will have one of the following six forms:
1 0 1 0 1 0 ⋆ ⋆ ⋆ ⋆ ⋆ ⋆
0 1 0 ⋆ 0 ⋆ 1 0 1 0 ⋆ ⋆
0 0
0 1
0 ⋆
0 1
0 ⋆
1
(11.7)
0
0 0 0 0 0 1 0 0 0 1 0 1
Note that there is only one subspace W for which M ends up with the first form. At the other extreme, if
the ground field k is infinite and you choose the space W uniformly at random (for basically any sensible
measure on Gr(k, kn )), then you will almost always end up with a matrix M of the last form.
285
2. For every p ∈ [n] ∼ |p|
k , there is a homeomorphism Ωp = k , where |p| = (p1 − 1) + (p2 − 2) + · · · + (pk − k) =
k+1
p1 + p2 + · · · + pk − 2 .
3. Define a partial order on [n]
k as follows: for p = {p1 < · · · < pk } and q = {q1 < · · · < qk }, set p ≥ q if
pi ≥ qi for every i. Then
p ≥ q =⇒ Ωp ⊇ Ωq . (11.8)
4. The poset [n]
k is isomorphic to the interval Yk,n in Young’s lattice.
5. Gr(k, kn ) is a compactification of the Schubert cell Ω(n−k+1,n−k+2,...,n) ∼
= kk(n−k) . In particular, dimk Gr(k, kn ) =
k(n − k).
For (2), the map Ωp → k|p| is given by reading off the ⋆s in the reduced column-echelon form of M (W ).
(For instance, let n = 4 and k = 2. Then the matrix representations in (11.7) give explicit diffeomorphisms
of the Schubert cells of Gr(k, kn ) to k0 , k1 , k2 , k2 , k3 , k4 respectively.) The number of ⋆s in the i-th column is
pi − i (pi − 1 entries above the pivot, minus i − 1 entries to the right of previous pivots), so the total number
of ⋆s is |p|.
For (3): This is best illustrated by an example. Consider the second matrix in (11.7):
1 0
0 z
M = 0 1
0 0
where I have replaced the entry labeled ⋆ by a parameter z. Here’s the trick: Multiply the second column
of this matrix by the scalar 1/z. Doing this doesn’t change the column span, i.e.,
1 0
0 1
colspace M = colspace 0 1/z .
0 0
which is the first matrix in (11.7). Therefore, the Schubert cell Ω1,2 is in the closure of the Schubert cell Ω1,3 .
In general, decrementing a single element of p corresponds to taking a limit of column spans in this way,
so the covering relations in the poset [n]
k give containment relations of the form (11.8).
Assertion (4) is purely combinatorial. The elements of Yk,n are partitions λ = (λ1 , . . . , λk ) such that n − k ≥
λ1 > · · · > λk ≥ 0. The desired poset isomorphism is p 7→ λp = (pk − k, pk−1 − (k − 1), . . . , p1 − 1). For
286
example, starting with (11.7)
1 0 1 0 1 0 ⋆ ⋆ ⋆ ⋆ ⋆ ⋆
0 1 0 ⋆ 0 ⋆ 1 0 1 0 ⋆ ⋆
Matrix
0
0 0 1 0 ⋆ 0 1 0 ⋆ 1 0
0 0 0 0 0 1 0 0 0 1 0 1
p 12 13 14 23 24 34
λp ∅
[n]
(5) now follows because p = (n − k + 1, n − k + 2, . . . , n) is the unique maximal element of k , and an easy
calculation shows that |p| = k(n − k).
This theorem amounts to a description of Gr(k, kn ) as a cell complex. (If you have not heard the term “cell
complex” before, now you know what it means: a topological space that is the disjoint union of cells —
that is, of homeomorphic copies of vector spaces — such that the closure of every cell is itself a union of
cells.) Furthermore, the poset isomorphism with Yk,n says that for every i, the number of cells of Gr(k, kn )
of dimension i is precisely the number of Ferrers diagrams with i blocks that fit inside the rectangle k n−k .
Combinatorially, we may write this equality as follows:
X X n
(# Schubert cells of dimension i) q i = #{λ ⊆ k n−k } q i = .
i i
k q
Example 11.4.3. If k = 1, then Gr(1, kn ) is the space of lines through the origin in kn ; that is, projective
space kP n−1 . As a cell complex, this has one cell of every dimension. For instance, the projective plane is
the union of three cells of dimensions 2, 1, and 0, i.e., a plane, a line and a point. In the standard geometric
picture, the 1-cell and 0-cell together form the “line at infinity”. Meanwhile, the interval Yk,n is a chain of
rank n − 1. Its rank-generating function is 1 + q + q 2 + · · · + q n−1 . (For k = C, double the dimensions of all
the cells, and substitute q 2 for q.) ◀
Remark 11.4.4. If k = C, then Gr(k, Cn ) is a cell complex with no odd-dimensional cells (because, topolog-
ically, the dimension of cells is measured over R). Therefore, readers who know some algebraic topology
(see, e.g., [Hat02, §2.2]) may observe that the cellular boundary maps are all zero (because each one has ei-
ther zero domain or zero range), so the cellular homology groups are exactly the chain groups. That is, the
Poincaré series of Gr(k, Cn ) is exactly the generating function for the dimensions of the cells. On the other
hand, If k = R, then the boundary maps need not be zero, and the homology can be more complicated.
Indeed, Gr(1, Rn ) = RP n−1 has torsion homology in odd dimensions.
Example 11.4.5. Let n = 4 and k = 2. Here is Yk,n :
287
These six partitions correspond to the six matrix-types in (11.7). The rank-generating function is
(1 − q 4 )(1 − q 3 )
4
= = 1 + q + 2q 2 + q 3 + q 4 .
2 q (1 − q 2 )(1 − q)
◀
Remark 11.4.6. What does all this have to do with enumerative geometry questions such as the Four-Lines
Problem? The answer (modulo technical details) is that the cohomology ring H ∗ (X) encodes intersections
of subvarieties3 of X: for every subvariety Z ⊆ Gr(k, kn ) of codimension i, there is a corresponding element
[Z] ∈ H i (X) (the “cohomology class of Z”) such that [Z ∪ Z ′ ] = [Z] + [Z ′ ] and [Z ∩ Z ′ ] = [Z][Z ′ ]. These
equalities hold only if Z and Z ′ are in general position with respect to each other (which has to be defined
precisely), but the consequence is that the Four-Lines Problem reduces to a computation in H ∗ (Gr(k, kn )):
find the cohomology class [Z] of the subvariety
Z = {W ∈ Gr(2, C4 ) : W meets some plane in C4 nontrivially}
and compare [Z]4 to the cohomology class [•] of a point. In fact, [Z]4 = 2[•]; this says that the answer to the
Four-Lines Problem is two, which is hardly obvious! To carry out this calculation, one needs to calculate
an explicit presentation of the ring H ∗ (Gr(k, kn )) as a quotient of a polynomial ring (which requires the
machinery of line bundles and Chern classes, but that’s another story) and then figure out how to express
the cohomology classes of Schubert cells with respect to that presentation. This is the theory of Schubert
polynomials.
There is a corresponding theory for the flag variety, which is the set F ℓ(n) of nested chains of vector spaces
F• = (0 = F0 ⊆ F1 ⊆ · · · ⊆ Fn = kn )
chains in the (infinite) lattice Ln (k). The flag variety is in fact a smooth manifold
or equivalently saturated
over k of dimension n2 . Like the Grassmannian, it has a decomposition into Schubert cells Xw , which are
indexed by permutations w ∈ Sn rather than partitions, as we now explain.
For every flag F• , we can find a vector space basis {v1 , . . . , vn } for kn such that Fk = k⟨v1 , . . . , vk ⟩ for all k,
and represent F• by the invertible matrix M ∈ G = GL(n, k) whose columns are v1 , . . . , vn . OTOH, any
ordered basis of the form
v1′ = b11 v1 , v2′ = b12 v1 + b22 v2 , . . . , vn′ = b1n v1 + b2n v2 + · · · + bnn vn ,
where bkk ̸= 0 for all k, defines the same flag. That is, a flag is a coset of B in G, where B is the subgroup
of invertible upper-triangular matrices (the Borel subgroup). Thus the flag variety can be (and often is)
regarded as the quotient G/B. This immediately implies that it is an irreducible algebraic variety (as G is
irreducible, and any image of an irreducible variety is irreducible). Moreover, it is smooth (e.g., because
every point looks like every other point, and so either all points are smooth or all points are singular and
the latter is impossible) and its dimension is (n − 1) + (n − 2) + · · · + 0 = n2 .
As in the case of the Grassmannian, there is a canonical representative for each coset of B, obtained by
Gaussian elimination, and reading off its pivot entries gives a decomposition
a
F ℓ(n) = Xw .
w∈Sn
3 If
you are more comfortable with differential geometry than algebraic geometry, feel free to think “submanifold” instead of “sub-
variety”.
288
Here the dimension of a Schubert cell Xw is the number of inversions of w, i.e.,
Recall that this is the rank function of the Bruhat and weak Bruhat orders on Sn . In fact, the (strong) Bruhat
order is the cell-closure partial order (analogous to (11.8)). It follows that the Poincaré polynomial of F ℓ(n)
is the rank-generating function of Bruhat order, namely
(1 + q)(1 + q + q 2 ) · · · (1 + q + · · · + q n−1 ).
More strongly, it can be shown that the cohomology ring H ∗ (F ℓ(n); Z) is the quotient of Z[x1 , . . . , xn ] by
the ideal generated by symmetric functions.
where ≤ means (strong) Bruhat order (see Ex. 1.2.13). These are much-studied objects in combinatorics;
for example, determining which Schubert varieties are singular turns out to to be a combinatorial question
involving the theory of pattern avoidance. Even more generally, instead of Sn , start with any finite Coxeter
group G (roughly, a group generated by elements of order two — think of them as reflections). Then G has a
combinatorially well-defined partial order also called the Bruhat order, and one can construct a G-analogue
of the flag variety: that is, a smooth manifold whose structure as a cell complex is given by Bruhat order
on G.
We now describe the calculation of the cohomology ring of F ℓ(n) using Chern classes. This is not intended
to be self-contained, and many facts will be presented as black boxes. For the full story, refer to, e.g., [BT82].
Definition 11.5.1. Let B and F be topological spaces. A bundle with base B and fiber F is a space E
together with a map π : E → B such that
1. If b ∈ B, then π −1 (b) ∼
= F ; and, more strongly,
2. Every b ∈ B has an open neighborhood U of b such that V := π −1 (U ) ∼
= U × F , and π|V is just
projection on the first coordinate.
Think of a bundle as a family of copies of F parameterized by B and varying continuously. The simplest
example of a bundle is a Cartesian product B × F with π(b, f ) = b; this is called a trivial bundle. Very often
the fiber is a vector space of dimension d, when we call the bundle a vector bundle of rank d; when d = 1
the bundle is a line bundle.
Frequently we require all these spaces to lie in a more structured category than that of topological spaces,
and we require the projection map to be a morphism in that category (e.g., manifolds with diffeomorphisms,
or varieties with algebraic maps).
Example 11.5.2. An example of a nontrivial bundle is a Möbius strip M , where B = S 1 is the central circle
and F = [0, 1] is a line segment. Indeed, a Möbius strip looks like a bunch of line segments parameterized
by a circle, and if U is any small interval in S 1 then the part of the bundle lying over U is just U × [0, 1].
However, the global structure of M is not the same as the cylinder S 1 × I. ◀
Example 11.5.3. Another important example is the tautological bundle on projective space Pd−1 k = Gr(1, kd ).
Recall that this is the space of lines ℓ through the origin in k . The tautological bundle T is the line bundle
d 4
defined by Tℓ = ℓ. That is, the fiber over a line is just the set of points on that line. ◀
4 The standard symbol for the tautological bundle is actually O(−1); let’s not get into why.
289
Let k be either R or C, and let us work in the category of closed compact manifolds over k. A vector bundle
of rank d is a bundle whose fiber is kd . (For example, the tautological bundle is a vector bundle of rank 1.)
Standard operations on vector spaces (direct sum, tensor product, dual, etc.) carry over to vector bundles,
defined fiberwise.
Let E be a rank-d vector bundle over M . Its projectivization P(E) is the bundle with fiber Pd−1
k defined by
P(E)m = P(Em ).
That is, a point in P(E) is given by a point m ∈ M and a line ℓ through the origin in Em ∼
= kd . In turn, P(E)
has a tautological line bundle L = L (E) whose fiber over (ℓ, m) is ℓ.
Associated with the bundle E are certain Chern classes ci (E) ∈ H 2i (M ) for every i, which measure “how
twisty E is.” (The 2 happens because we are talking about a complex manifold.) I will not define these
classes precisely (see [BT82]), but instead will treat them as a black box that lets us calculate cohomology.
The Chern classes have the following properties:
1. c0 (E) = 1 by convention.
2. ci (E) = 0 for i > rank E.
3. If E is trivial then ci (E) = 0 for i > 0.
4. If 0 → E ′ → E → E ′′ → 0 is an exact sequence of M -bundles, then c(E) = c(E ′ )c(E ′′ ), where c(E) =
P
i ci (E) (the “total Chern class”).
5. For a line bundle L, c1 (L∗ ) = −c1 (L).
Here is the main formula, which expresses the cohomology ring of a bundle as a module over the cohomol-
ogy of its base.
where x = c1 (L ).
Example 11.5.4 (Projective space). Pd−1 C is the projectivization of the trivial rank-d bundle over M = {•}.
Of course H ∗ (M ; Z) = Z, so H ∗ (Pd−1 C; Z) = Z[x]/⟨xd ⟩. ◀
Example 11.5.5 (The flag variety F ℓ(3)). Let M = P2 = Gr(1, C3 ). Define a bundle E 2 by
Eℓ2 = C3 /ℓ.
Then E 2 has rank 2, and P(E 2 ) is just the flag variety F ℓ(3), because specifying a line in C3 /ℓ is the same
thing as specifying a plane in C3 containing ℓ. Let L = L (E 2 ). For each ℓ ∈ M we have an exact sequence
0 → ℓ → C3 → C3 /ℓ → 0, which gives rise to a short exact sequence of bundles
0 → O → C3 → E 2 → 0
where O is the tautological bundle on M , with c1 (O) = x (the generator of H ∗ (M )). The rules for Chern
classes then so the rules for Chern classes tell us that
(1 + x)(1 + c1 (E 2 ) + c2 (E 2 )) = 1
x + c1 (E 2 ) = 0, xc1 (E 2 ) + c2 (E 2 ) = 0
290
In fact this ring is isomorphic to
X1 = P(E0 ). Let E1 be the rank-(n − 1) bundle whose fiber over a line E1 is Cn /E1 .
X2 = P(E1 ). This is the partial flag variety of flags E• : 0 = E0 ⊆ E1 ⊆ E2 . Let E2 be the rank-(n − 2)
bundle whose fiber over E• is Cn /E2 .
We end up with generators x1 , . . . , xn , one for the tautological bundle of each Ei . The relations turn out to
be the symmetric functions on them. That is.
H ∗ (F ℓ(n)) ∼
= Q[x1 , . . . , xn ]/⟨e1 , e2 , . . . , en ⟩
The Poincare polynomial of the flag variety (i.e., the Hilbert series of its cohomology ring) can be worked
out explicitly. Modulo the elementary symmetric functions, every polynomial can be written as a sum of
monomials of the form
xa1 1 xa2 2 · · · xann
where ai < i for all i. Therefore,
X
Poin(F ℓ(n), q) = q k dimQ H 2i (F ℓ(n)) = (1)(1 + q)(1 + q + q 2 ) · · · (1 + q + · · · + q n−1 ) = [q]n !.
k
where Sn is the symmetric group on n letters and inv(w) is the number of inversions:
In fact the flag variety has a natural cell decomposition into Schubert cells. Given any flag
E• : 0 = E0 ⊆ E1 ⊆ · · · ⊆ En = Cn
construct a n × n matrix [v1 | · · · |vn ] in which the first k columns are a basis of Ek , for every k. We can
canonicalize the matrix as follows:
291
• Scale the first column so that its bottom nonzero entry is 1. Say this occurs in row w1 .
• Add an appropriate multiple of v1 to each of v2 , . . . , vn so as to kill off the entry in row w1 . Note that
this does not change the flag.
• Scale the second column so that its bottom nonzero entry is 1. Say this occurs in row w2 . Note that
w2 ̸= w1 .
• Add an appropriate multiple of v2 to each of v3 , . . . , vn so as to kill off the entry in row w1 .
• Repeat.
We end up with a matrix that includes a “pivot” 1 in each row and column, with zeroes below and to the
right of every 1. The pivots define a permutation w ∈ Sn . For example, if w = 4132 then the matrix will
have the form
∗ 1 0 0
∗ 0 ∗ 1
∗ 0 1 0 .
1 0 0 0
◦
The set X3142 of all matrices of this type is a subspace of F ℓ(4) that is in fact isomorphic to C3 — the stars
are affine coordinates. Thus we obtain a decomposition into Schubert cells
a
◦
F ℓ(n) = Xw
w∈Sn
and moreover the stars correspond precisely to inversions of w. This gives the Poincaré polynomial.
The closure of a Schubert cell is called a Schubert variety. The cohomology classes of Schubert varieties
are also a vector space basis for H ∗ (F ℓ(n)), and there is a whole theory of how to translate between the
“algebraic” basis (coming from line bundles) and the “geometric” basis (Schubert varieties).
11.6 Exercises
292
For (11.3b), the conservation constraints (11.2), together with (11.3a), imply that
X
f + (S) − f − (S) = (f + (v) − f − (v)) = f + (s) − f − (s) = |f | = f − (t) − f + (t) = f − (T ) − f + (T ).
v∈S
Problem 11.2. Let G(V, E) be a graph. A matching on G is a collection of edges no two of which share an
endpoint. A vertex cover is a set of vertices that include at least one endpoint of each edge of G. Let µ(G)
denote the size of a maximum matching, and let β(G) denote the size of a minimum vertex cover.
(a) (Warmup) Show that µ(G) ≤ β(G) for every graph G. Exhibit a graph for which the inequality is
strict.
(b) The König-Egerváry Theorem asserts that µ(G) = β(G) whenever G is bipartite, i.e., the vertices of G
can be partitioned as X ∪ Y so that every edge has one endpoint in each of X, Y . Derive the König-
Egerváry Theorem as a consequence of the Max-Flow/Min-Cut Theorem.
(c) Prove that the König-Egerváry Theorem and Dilworth’s Theorem imply each other.
Solution: To be written.
Polyá theory
Problem 11.3. Let n ≥P2 and for σk ∈ Sn , let f (σ) denote the number of fixed points. Prove that for every
1
k ≥ 1, the number n! σ∈Sn f (σ) is an integer.
Solution: [D. Grinberg] Apply Burnside’s Theorem (Prop. 11.3.2) to the component-wise action of Sn on
[n]k .
293
Appendix: Catalan Numbers
A Dyck path of size n is a path from (0, 0) to (2n, 0) in R2 consisting of n up-steps and n down-steps that
stays (weakly) above the x-axis.
We can denote Dyck paths efficiently by a list of U’s and D’s; the path P shown above is UUDUUDDD. Each
up-step can be thought of as a left parenthesis, and each down-step as a right parenthesis, so we could also
write P = (()(())). The requirement of staying above the x-axis then says that each right parenthesis must
close a previous left parenthesis.
Proposition 11.6.1. The number of Dyck paths of size n is the Catalan number Cn .
Sketch of proof. The proof is an illustration of the Sheep Principle (“in order to count the sheep in a flock,
count the legs and divide by four”). Consider the family L of all lattice paths from (0, 0) to (2n + 1, −1)
consisting of n up-steps and n + 1 down-steps (with no restrictions); evidently |L| = 2n+1n .
Consider the action of the cyclic group Z2n+1 on L by cyclic rotation. First, the orbits all have size 2n + 1.
(There is no way that a nontrivial element of Z2n+1 can fix the locations of the up-steps, essentially because
gcd(2n + 1, n) = 1 — details left to the reader.) Second, each orbit contains exactly one augmented Dyck path,
294
i.e., a Dyck path followed by a down-step. (Of all the lowest points in a path, find the leftmost one and call
it z. Rotate so that the last step is the down-step that lands at z.)
Figure 11.6: Rotating the lattice path UDDUDD|UDU to obtain the augmented Dyck path UDU|UDDUDD.
Every (augmented) Dyck path arises in this way, so we have a bijection. The orbits are sheep and each
sheep has 2n + 1 legs, so the number of Dyck paths is
1 2n + 1 (2n + 1)! (2n)! (2n)! 1 2n
= = = = .
2n + 1 n (2n + 1) (n + 1)! n! (n + 1)! n! (n + 1) n! n! n+1 n
To show that a class of combinatorial objects is enumerated by the Catalan numbers, one can now find a
bijection to Dyck paths. A few of the most commonly encountered interpretations of Cn are:
Others will be encountered in the course of these notes. For details, see [Sta99] or [Sta15] Another core
feature of the Catalan numbers is that they satisfy the following recurrence:
n−1
X
Cn = Cn−1 + Ck−1 Cn−k for n ≥ 1. (11.10)
k=1
This equation can be checked by a banal induction argument, but it is also worthwhile seeing the combina-
torial reason for it. Call a Dyck path of size n primitive if it stays strictly above the x-axis for 0 < x < 2n.
If a path P is primitive, then it is of the form UP ′ D for some Dyck path P ′ of size n − 1 (not necessarily
primitive); this accounts for the Cn−1 term in the Catalan recurrence. Otherwise, let (2k, 0) be the smallest
positive x-intercept, so that 1 ≤ k ≤ n − 1. The part of the path from (0, 0) to (2k, 0) is a primitive Dyck
path of size k, and the part from (2k, 0) to (2n, 0) is a Dyck path of size n − k, not necessarily primitive.
295
296
Notational Index
Basics
◀ End of an example
[n] {1, . . . , n}
N nonnegative integers 0, 1, 2, . . . \mathbb{N}
N>0 positive integers 1, 2, . . . \mathbb{P}
2S power set of a set S (or the associated poset)
|S| or #S cardinality of set S
∪· disjoint union \cupdot (requires [Link])
△ symmetric difference A△B = (A ∪ B) \ (A ∩ B) \triangle
Sn symmetric group on n letters \mathfrak{S}_n
S
k set of k-element subsets of a set S \binom{S}{k}
Cn Catalan numbers
k⟨v1 , . . . , vn ⟩ k-vector space with basis {v1 , . . . , vn } \fld\langle ... \rangle
Posets
Lattices
Hyperplane Arrangements
298
Representation Theory
Symmetric Functions
299
Combinatorial Algebraic Varieties
µ product
∆ coproduct
u unit
ε counit
S antipode
300
Bibliography
[AA17] Marcelo Aguiar and Federico Ardila, Hopf monoids and generalized permutahedra, preprint,
arXiv:1709.07504, 2017. 169, 271
[ABS06] Marcelo Aguiar, Nantel Bergeron, and Frank Sottile, Combinatorial Hopf algebras and generalized
Dehn-Sommerville relations, Compos. Math. 142 (2006), no. 1, 1–30. MR 2196760 271
[AM10] Marcelo Aguiar and Swapneel Mahajan, Monoidal functors, species and Hopf algebras, CRM
Monograph Series, vol. 29, American Mathematical Society, Providence, RI, 2010, With
forewords by Kenneth Brown and Stephen Chase and André Joyal. MR 2724388 271
[Ath96] Christos A. Athanasiadis, Characteristic polynomials of subspace arrangements and finite fields,
Adv. Math. 122 (1996), no. 2, 193–233. MR 1409420 (97k:52012) 124, 125
[BB05] Anders Björner and Francesco Brenti, Combinatorics of Coxeter Groups, Graduate Texts in
Mathematics, vol. 231, Springer, New York, 2005. MR 2133266 (2006d:05001) 17, 32
[BF01] Thomas Britz and Sergey Fomin, Finite posets and Ferrers shapes, Adv. Math. 158 (2001), no. 1,
86–127. MR 1814900 280
[BH93] Winfried Bruns and Jürgen Herzog, Cohen-Macaulay Rings, Cambridge Studies in Advanced
Mathematics, vol. 39, Cambridge University Press, Cambridge, 1993. MR 1251956 (95h:13020)
142, 147
[Bjö94] Anders Björner, Subspace arrangements, First European Congress of Mathematics, Vol. I (Paris,
1992), Progr. Math., vol. 119, Birkhäuser, Basel, 1994, pp. 321–370. MR 1341828 131
[BLVS+ 99] Anders Björner, Michel Las Vergnas, Bernd Sturmfels, Neil White, and Günter M. Ziegler,
Oriented matroids, second ed., Encyclopedia of Mathematics and its Applications, vol. 46,
Cambridge University Press, Cambridge, 1999. MR 1744046 134
[BM71] H. Bruggesser and P. Mani, Shellable decompositions of cells and spheres, Math. Scand. 29 (1971),
197–205 (1972). MR 0328944 (48 #7286) 166
[BNP98] Marilena Barnabei, Giorgio Nicoletti, and Luigi Pezzoli, Matroids on partially ordered sets, Adv.
in Appl. Math. 21 (1998), no. 1, 78–112. MR 1623325 84
[BO92] Thomas Brylawski and James Oxley, The Tutte polynomial and its applications, Matroid
applications, Encyclopedia Math. Appl., vol. 40, Cambridge Univ. Press, Cambridge, 1992,
pp. 123–225. MR 1165543 (93k:05060) 111
[Bol98] Béla Bollobás, Modern Graph Theory, Graduate Texts in Mathematics, vol. 184, Springer-Verlag,
New York, 1998. MR 1633290 (99h:05001) 98
301
[BR15] Matthias Beck and Sinai Robins, Computing the continuous discretely, second ed., Undergraduate
Texts in Mathematics, Springer, New York, 2015., Full text available free from
[Link] MR 3410115 104, 169
[Bri73] Egbert Brieskorn, Sur les groupes de tresses [d’après V. I. Arnol′ d], Séminaire Bourbaki, 24ème
année (1971/1972), Exp. No. 401, Springer, Berlin, 1973, pp. 21–44. Lecture Notes in Math., Vol.
317. 131
[BT82] Raoul Bott and Loring W. Tu, Differential Forms in Algebraic Topology, Graduate Texts in
Mathematics, vol. 82, Springer-Verlag, New York-Berlin, 1982. MR 658304 (83i:57016) 289, 290
[CR70] Henry H. Crapo and Gian-Carlo Rota, On the Foundations of Combinatorial Theory: Combinatorial
Geometries, preliminary ed., The M.I.T. Press, Cambridge, Mass.-London, 1970. MR 0290980 124
[Dir61] G.A. Dirac, On rigid circuit graphs, Abh. Math. Sem. Univ. Hamburg 25 (1961), 71–76. MR
0130190 (24 #A57) 130
[Edm70] Jack Edmonds, Submodular functions, matroids, and certain polyhedra, Combinatorial Structures
and their Applications (Proc. Calgary Internat. Conf., Calgary, Alta., 1969), Gordon and
Breach, New York, 1970, pp. 69–87. MR 0270945 169
[FS05] Eva Maria Feichtner and Bernd Sturmfels, Matroid polytopes, nested sets and Bergman fans, Port.
Math. (N.S.) 62 (2005), no. 4, 437–468. MR 2191630 169
[Ful97] William Fulton, Young Tableaux, London Mathematical Society Student Texts, vol. 35,
Cambridge University Press, Cambridge, 1997. 235, 253, 284
[GGMS87] I. M. Gel’fand, R. M. Goresky, R. D. MacPherson, and V. V. Serganova, Combinatorial geometries,
convex polyhedra, and Schubert cells, Adv. in Math. 63 (1987), no. 3, 301–316. MR 877789 169
[GK76] Curtis Greene and Daniel J. Kleitman, The structure of Sperner k-families, J. Combin. Theory Ser.
A 20 (1976), no. 1, 41–68. 280
[GNW79] Curtis Greene, Albert Nijenhuis, and Herbert S. Wilf, A probabilistic proof of a formula for the
number of Young tableaux of a given shape, Adv. in Math. 31 (1979), no. 1, 104–109. MR 521470 248
[Gre76] Curtis Greene, Some partitions associated with a partially ordered set, J. Combin. Theory Ser. A 20
(1976), no. 1, 69–79. 280
[Grü03] Branko Grünbaum, Convex Polytopes, second ed., Graduate Texts in Mathematics, vol. 221,
Springer-Verlag, New York, 2003, Prepared and with a preface by Volker Kaibel, Victor Klee
and Günter M. Ziegler. 75, 160
[GSS93] Jack Graver, Brigitte Servatius, and Herman Servatius, Combinatorial Rigidity, Graduate Studies
in Mathematics, vol. 2, American Mathematical Society, Providence, RI, 1993. MR 1251062 63
[Hat02] Allen Hatcher, Algebraic Topology, Cambridge University Press, Cambridge, 2002. MR 1867354
(2002k:55001) 142, 144, 146, 287
[HS15] Joshua Hallam and Bruce Sagan, Factoring the characteristic polynomial of a lattice, J. Combin.
Theory Ser. A 136 (2015), 39–63. MR 3383266 130
[Hum90] James E. Humphreys, Reflection Groups and Coxeter Groups, Cambridge Studies in Advanced
Mathematics, vol. 29, Cambridge University Press, Cambridge, 1990. MR 1066460 (92h:20002)
17
[Hum11] Brandon Humpert, Polynomials Associated with Graph Coloring and Orientations, Ph.D. thesis,
University of Kansas, 2011. 270
302
[Kli] Caroline J. Klivans, A quasisymmetric function for generalized permutahedra, Talk slides,
[Link] 271
[LS00] Shu-Chung Liu and Bruce E. Sagan, Left-modular elements of lattices, J. Combin. Theory Ser. A 91
(2000), no. 1-2, 369–385. MR 1780030 57
[MR05] Jeremy L. Martin and Victor Reiner, Cyclotomic and simplicial matroids, Israel J. Math. 150 (2005),
229–240. MR 2255809 91
[MS05] Ezra Miller and Bernd Sturmfels, Combinatorial Commutative Algebra, Graduate Texts in
Mathematics, vol. 227, Springer-Verlag, New York, 2005. MR 2110098 (2006d:13001) 142
[OS80] Peter Orlik and Louis Solomon, Combinatorics and topology of complements of hyperplanes, Invent.
Math. 56 (1980), no. 2, 167–189. 131
[OT92] Peter Orlik and Hiroaki Terao, Arrangements of Hyperplanes, Grundlehren der Mathematischen
Wissenschaften [Fundamental Principles of Mathematical Sciences], vol. 300, Springer-Verlag,
Berlin, 1992. 112
[Oxl92] James G. Oxley, Matroid Theory, Oxford Science Publications, The Clarendon Press Oxford
University Press, New York, 1992. 63, 76, 90
[Per16] Xander Perrott, Existence of projective planes, arXiv:1603.05333 (retrieved 8/26/18), 2016. 75
[Pos09] Alexander Postnikov, Permutohedra, associahedra, and beyond, Int. Math. Res. Not. (2009), no. 6,
1026–1106. MR 2487491 169
[PRW08] Alex Postnikov, Victor Reiner, and Lauren Williams, Faces of generalized permutohedra, Doc.
Math. 13 (2008), 207–273. MR 2520477 134, 169
[Rei] Victor Reiner, Lectures on matroids and oriented matroids, Lecture notes for Algebraic
Combinatorics in Europe (ACE) Summer School in Vienna, July 2005; available at
[Link] 134, 138
[RGZ97] Jürgen Richter-Gebert and Günter M. Ziegler, Oriented matroids, Handbook of discrete and
computational geometry, CRC Press Ser. Discrete Math. Appl., CRC, Boca Raton, FL, 1997,
pp. 111–132. MR 1730162 134, 135
[Ryb11] G.L. Rybnikov, On the fundamental group of the complement of a complex hyperplane arrangement,
Funktsional. Anal. i Prilozhen. 45 (2011), no. 2, 71–85. 131
[S+ 14] W. A. Stein et al., Sage Mathematics Software (Version 6.4.1), The Sage Development Team, 2014,
[Link] 113
[Sag01] Bruce E. Sagan, The Symmetric Group, second ed., Graduate Texts in Mathematics, vol. 203,
Springer-Verlag, New York, 2001. 223
[Sch13] Alexander Schrijver, A Course in Combinatorial Optimization, Available online at
[Link] (retrieved 1/21/15), 2013. 160
[Sta95] Richard P. Stanley, A symmetric function generalization of the chromatic polynomial of a graph, Adv.
Math. 111 (1995), no. 1, 166–194. MR 1317387 262
[Sta96] , Combinatorics and Commutative Algebra, second ed., Progress in Mathematics, vol. 41,
Birkhäuser Boston Inc., Boston, MA, 1996. MR 1453579 (98h:05001) 142
303
[Sta99] , Enumerative Combinatorics. Vol. 2, Cambridge Studies in Advanced Mathematics,
vol. 62, Cambridge University Press, Cambridge, 1999, With a foreword by Gian-Carlo Rota
and appendix 1 by Sergey Fomin. MR 1676282 (2000k:05026) 216, 221, 229, 235, 240, 241, 294,
295
[Sta07] , An introduction to hyperplane arrangements, Geometric combinatorics, IAS/Park City
Math. Ser., vol. 13, Amer. Math. Soc., Providence, RI, 2007, Also available online:
[Link] pp. 389–496. 112, 126,
129, 140
[Sta12] , Enumerative Combinatorics. Volume 1, second ed., Cambridge Studies in Advanced
Mathematics, vol. 49, Cambridge University Press, Cambridge, 2012. MR 2868112 46, 169
[Sta15] , Catalan Numbers, Cambridge University Press, New York, 2015. MR 3467982 294, 295
[Tak71] Mitsuhiro Takeuchi, Free Hopf algebras generated by coalgebras, J. Math. Soc. Japan 23 (1971),
561–582. MR 0292876 269
[Tut54] W. T. Tutte, A contribution to the theory of chromatic polynomials, Canad. J. Math. 6 (1954), 80–91.
MR 61366 100
[Wes96] Douglas B. West, Introduction to Graph Theory, Prentice Hall, Inc., Upper Saddle River, NJ, 1996.
MR 1367739 (96i:05001) 130
[Whi32a] Hassler Whitney, The coloring of graphs, Ann. of Math. (2) 33 (1932), no. 4, 688–718. MR 1503085
121
[Whi32b] , A logical expansion in mathematics, Bull. Amer. Math. Soc. 38 (1932), no. 8, 572–579. MR
1562461 121
[Zas75] Thomas Zaslavsky, Facing up to arrangements: face-count formulas for partitions of space by
hyperplanes, Mem. Amer. Math. Soc. 1 (1975), no. issue 1, 154, vii+102. MR 0357135 118
[Zie95] Günter M. Ziegler, Lectures on polytopes, Graduate Texts in Mathematics, vol. 152,
Springer-Verlag, New York, 1995. MR 1311028 160, 162, 168
304
A “pointillist” picture of the essentialized braid arrangement ess(Br4 ), produced by a computer glitch.
305