Introduction to Grassmann Manifolds
Introduction to Grassmann Manifolds
where ai = (aji )j∈{1,...,n} ∈ Rn , i ∈ {1, . . . , k}. We endow Rn×k with the norm induced
by the Frobenius inner product, i.e.
q sX
>
>
kAk = tr(A A) = |a1 |2 · · · |ak |2 = a2ij . (1)
2
i,j
∗
This manuscript was written while doing my PhD at Fachrichtung Mathematik, Technische Univer-
sität Dresden, 01062 Dresden, Germany. Funding by ESF grant “Landesinnovationspromotion” No.
080942988 is gratefully acknowledged.
†
E-mail: kadaniel@[Link]
1
We define the (non-compact) Stiefel manifold as
n o
St(k, Rn ) := A ∈ Rn×k ; rk(A) = k
n o
= A = a1 · · · ak ∈ Rn×k ; a1 , . . . , ak are linearly independent .
St(k, Rn ) is often referred to as the set of all k-frames (of linearly independent vectors)
in Rn . Note that St(k, Rn ) is an open subset of Rn×k (as the preimage of the open
set R \ {0} under the continuous map A 7→ det(A> A)). We define the compact Stiefel
manifold as n o
St∗ (k, Rn ) := A ∈ St(k, Rn ); A> A = Ik .
√
Clearly, St∗ (k, Rn ) is a bounded (by k) and closed (as the preimage of the closed set
{Ik } under the continuous map A 7→ A> A) and hence compact subset of St(k, Rn ) and
St∗ (k, Rn ) ,→ St(k, Rn ). It is often referred to as the set of all orthonormal k-frames of
linearly independent vectors in Rn .
We consider the map
π : St(k, Rn ) → Gr(k, Rn ),
A = a1 · · · ak 7→ span {a1 , . . . , ak } , (2)
and endow Gr(k, Rn ) with the final topology with respect to π, which we call the Grass-
mann topology, i.e. U ⊆ Gr(k, Rn ) is open if and only if π −1 [U ] is open in St(k, Rn ).
Let f ∈ L(Rk , Rn ) be injective and A ∈ St(k, Rn ) be the matrix representation with
respect to the canonical bases in Rk and Rn . Then π(A) can be equivalently considered
as f [Rk ], i.e. the image of f .
1.1 Lemma. Consider Gr(k, Rn ) with the Grassmann topology.
(i) The map π is surjective and for any A ∈ St(k, Rn ) we have
which can be endowed with the structure of a quotient manifold; see [2, Section 3.4].
However, we define a topological and differentiable structure on Gr(k, Rn ) directly.
2
We consider the restriction of π to St∗ (k, Rn ), i.e. π := π St∗ (k,Rn ) , and the “Gram-
Schmidt orthonormalization map”, i.e. GS(A) is the matrix obtained by the application
of the Gram-Schmidt orthonormalization method to the columns a1 , . . . , an of A. It is
well-known that the Gram-Schmidt method defines a continuous map. Thus, we have
π ◦ GS = π. Clearly, π is surjective and continuous.
Next, we show a characterization of the continuity of functions f defined on a topo-
logical space endowed with the final topology with respect to some other function and
mapping to another topological space.
1.2 Lemma. Consider Gr(k, Rn ) with the Grassmann topology. Let Ω be a topological
space and
f : Gr(k, Rn ) → Ω.
Then the following statements are equivalent:
Proof. (a) ⇔ (b) follows from the definition of final topology with respect to π.
(b) ⇒ (c) is obvious since π is the restriction of π.
(c) ⇒ (b) follows from the representation f ◦ π = f ◦ π ◦ GS and the continuity of
GS.
where ΠX denotes the orthogonal projection on a subspace X and k·k denotes the op-
erator norm. Θ is called the gap metric on Gr(k, Rn ).
3
In the literature, Θ(L, M ) is also referred to as the opening between the subspaces L
and M and was originally introduced in [11]. By the Projection Theorem for Hilbert
spaces, there is a one-to-one correspondence between (closed) subspaces U and orthog-
onal projections ΠU with im ΠU = U and ker ΠU = U ⊥ . Therefore, it is easy to see that
Θ is indeed a metric on Gr(k, Rn ). Next we show some further properties of the gap
metric.
|x − ΠM x| = |(ΠL − ΠM ) x| ≤ kΠL − ΠM k .
Therefore,
sup d(x, M ) ≤ kΠL − ΠM k .
x∈S∩L
Analogously we have
sup d(x, L) ≤ kΠL − ΠM k
x∈S∩M
and we obtain
max sup d(x, M ), sup d(x, L) ≤ kΠL − ΠM k = Θ(L, M ).
x∈S∩L x∈S∩M
4
We estimate
and hence
|ΠM (I − ΠL )x| ≤ ρM |(I − ΠL )x| . (8)
On the other hand, using
ΠM − ΠL = ΠM (I − ΠL ) − (I − ΠM )ΠL
Thus, we obtain
kΠL − ΠM k ≤ max {ρM , ρL } ,
which finishes the proof of (a).
(ii) The equality Θ(L, M ) = Θ(L⊥ , M ⊥ ) follows directly from the definition, since
ΠL⊥ = I − ΠL . The estimate Θ(L, M ) ≤ 1 holds by the observation that for x ∈ S ∩ L
we have d(x, M ) = |ΠM ⊥ x| = |x| − |ΠM x| ≤ |x| = 1.
(iii) First note that the second equivalence is clear by the definition of decompositions.
The first one follows from the following observation.
1.6 Claim. Let M ∈ Gr(k, Rn ) and x ∈ S. Then x ∈ M ⊥ if and only if d(x, M ) = 1.
Proof of claim. The proof relies on the relation |x|2 = |ΠM x|2 + |ΠM ⊥ x|2 . From the
relation we directly read off that
Conversely, from d(x, M ) = |ΠM ⊥ x|2 = |x|2 = 1 follows |ΠM x| = 0 and hence x ∈
M ⊥.
5
Due to the continuity of the norm and the compactness of S ∩ L the value of Θ(L, M )
is attained, so from Θ(L, M ) < 1 follows L ∩ M ⊥ = L⊥ ∩ M = {0}. On the other hand,
from Θ(L, M ) = 1, which we (w.l.o.g.) assume to be attained by supx∈S∩L d(x, M ) = 1,
follows that there exists some z ∈ S ∩ L such that d(z, M ) = 1 and by the claim z ∈ M ⊥
holds.
Our next aims are to prove that the gap metric on Gr(k, Rn ) is equivalent to the Haus-
dorff metric on the sections of two subspaces with the unit sphere (called the spherical
gap), and second to show the equivalence of the Grassmann topology to the topology
induced by the gap metric/Hausdorff metric.
1.7 Definition (Spherical gap metric, cf. also [10]). We define
Θ(L, M ) ≤ Θ(L,
e M ) ≤ 2Θ(L, M ).
Proof. We follow [10, p. 199]. The first inequality follows trivially from d(x, M ) ≤
d(x, M ∩ S), x ∈ Rn . To show the second inequality it suffices to prove d(x, M ∩
S) ≤ 2d(x, M ) for x ∈ S. By the definition of infimum, for any ε ∈ R>0 there exists
y
y ∈ M \ {0} such that |x − y| < d(x, M ) + ε. Define y0 := |y| ∈ S ∩ M . Then the
estimate d(x, M ∩ S) ≤ |x − y0 | ≤ |x − y| + |y − y0 | holds. Furthermore, since y and y0
are parallel, we have
y
|y − y0 | = y − = ||y| − 1| = ||y| − |x|| ≤ |y − x| .
|y|
In summary, we have d(x, M ∩ S) < 2d(x, M ) + 2ε, and since ε is arbitrary, the prove is
finished.
1.10 Remark. For a sequence (Xn )n∈N ∈ (Gr(k, Rn ))N converging to X ∈ Gr(k, Rn )
n→∞
with respect to Θ, i.e. kΠXn − ΠX k −−−→ 0, we have by the definition of Hausdorff
metric that for any x ∈ X there exists a sequence (xn )n∈N ∈ (S)N with xn ∈ Xn for
n→∞
each n ∈ N such that xn −−−→ x, and hence with respect to any norm.
6
1.11 Remark. When Θ is introduced by Eq. (6) in the set of subspaces of general Banach
spaces, Θ does not necessarily satisfy the triangle inequality and can therefore not be
used to define a metric/topology. However, in Hilbert space one can show equality to the
expression given in Eq. (5), which we have used for the definition and which obviously
satisfies the conditions for a metric. On the other hand, Θ
e defines a proper metric even in
the general Banach space case; see, for instance, [10, IV.2.1] and the references therein.
Next we address our second aim, to show that the Grassmann topology coincides with
the topology induced by the gap metric Θ (or equivalently Θ). e To this end, we need to
prove that the identity map from the topological space Gr(k, Rn ) with the Grassmann
topology to the topological space Gr(k, Rn ) with the gap topology is continuous together
with its inverse. Since Gr(k, Rn ) with the Grassmann topology is compact and the
metric space Gr(k, Rn ) with the metric Θ is Hausdorff, the continuity of the inverse
follows directly from the continuity of the identity map. Furthermore, by Lemma 1.2 it
suffices to show that π : St∗ (k, Rn ) → (Gr(k, Rn ), Θ) is continuous.
1.12 Proposition. The Grassmann topology in Gr(k, Rn ) coincides with the topologies
induced by Θ and Θ,
e respectively.
7
2 Differentiable Structure
This section merges ideas of [6, 1]. To introduce the differentiable structure, we make
use of local affine cross sections. Let A ∈ St(k, Rn ) and define
n o n o
SA := B ∈ St(k, Rn ); A> (B − A) = 0 = B ∈ St(k, Rn ); A> B = A> A
n o
⊆ B ∈ St(k, Rn ); det(A> B) 6= 0 =: TA ⊂ St(k, Rn ),
orthogonal to the fiber A[GL(k, R)] crossing through A. For B ∈ St(k, Rn ) the equiva-
lence class B[GL(k, R)] = π −1 [π(B)] intersects the cross section SA if and only if B ∈ TA ,
and then in B(A> B)−1 A> A. This can be seen by plugging BP with P ∈ GL(k, R) in
the definition of SA :
Since for each A ∈ St(k, Rn ) we have A ∈ TA and TA is open, (TA )A∈St(k,Rn ) is an open
covering of St(k, Rn ). On the other hand, for B ∈ St(k, Rn ) the set
{A ∈ St(k, Rn ); B ∈ TA } = TB
be the set of subspaces whose representing fibers B[GL(k, R)] intersect the cross section
SA . We call the mapping
(iv) For each A ∈ St(k, Rn ) one has that σA is continuous, π ◦σA = idUA and σAP (L) =
σA (L)P for P ∈ GL(k, R) and L ∈ UA .
8
Proof. (i) The covering property is clear by the surjectivity of π and the fact that
(TA )A∈St(k,Rn ) is an open covering of St(k, Rn ). The fact that for each A ∈ St(k, Rn )
the set UA is open follows from the fact that both TA andS π are open.
(ii) We have by definition that π −1 [UA ] = π −1 [π[TA ]] = P ∈GL(k,R) [TA ]P = TA , since
TA is invariant under right-multiplication with elements from GL(k, R).
(iii) This is clear by the continuous embedding of TA in St(k, Rn ) via the identity.
(iv) In order to prove continuity of σA , by Lemma 1.2 it is equivalent to show the
continuity of σA ◦π : TA ⊂ St(k, Rn ) → SA , which is obvious by the representation in Eq.
(9). To see that π ◦ σA = id observe that for B ∈ TA we have (A> B)−1 A> A ∈ GL(k, R).
Let B ∈ TA and L := π(B) ∈ UA , then π(σA (L)) = π(B(A> B)−1 A> A) = π(B) = L.
We show the homogeneity property by calculating
σAP (π(B)) = B((AP )> B)−1 (AP )> AP = B(P > A> B)−1 (AP )> AP
= B(A> B)−1 (P > )−1 P > A> AP = B(A> B)−1 A> AP
= σA (π(B))P.
(v) This is clear with the fact that WB is open and with the representation in Eq.
(9).
and secondly
−1
> −1 > Σ
V > V ΣΣV > V Σ 0(n−k)×k U >
A(A A) A =U
0(n−k)×k
Σ
Σ−1 Σ−1 Σ 0(n−k)×k U >
=U
0(n−k)×k
9
Ik
Ik 0(n−k)×k U >
=U
0(n−k)×k
Ik 0k×(n−k)
=U U>
0(n−k)×k 0(n−k)×(n−k)
0k×k 0k×(n−k)
= U In − U>
0(n−k)×k In−k
= In − A⊥ A>
⊥. (11)
Note that (A> A)−1 A> is the Moore-Penrose inverse to A since A has full rank. It is well-
known that for B ∈ St∗ (k, Rn ) the matrix BB > represents the orthogonal projection
onto π(B). Thus, we derived in Eq. (11) that A(A> A)−1 A> represents the orthogonal
projection onto π(A) and equivalently In − A(A> A)−1 A> represents the orthogonal
projection onto π(A)⊥ .
With respect to A ∈ St(k, Rn ) we define the following family of functions
ϕA : UA → R(n−k)×k ,
(12)
L = π(B) 7→ A> > > −1 >
⊥ σA (L) = A⊥ B(A B) A A.
10
(c) Clearly, ϕA and pA are continuous for every A ∈ St(k, Rn ), i.e. ϕA is a homeo-
morphism. For A ∈ St(k, Rn ) the domain of ϕA is open, thus UA ∩ UB is open and
consequently ϕB [UA ∩ UB ] = p−1
B [UA ∩ UB ] ⊆ R
(n−k)×k is open due to continuity of p
B
n
for any B ∈ St(k, R ).
(d) Let A, B ∈ St(k, Rn ) be such that UA ∩ UB 6= ∅. Then for K ∈ ϕA [UB ] we have
>
ϕB ◦ pA (K) = B⊥ (A + A⊥ K)(B > (A + A⊥ K))−1 B > B
−1
> >
= B⊥ A + B⊥ A⊥ K B > A + B > A⊥ K B > B,
2.3 Corollary. Let Gr(k, Rn ) be endowed with the differentiable structure from Theo-
rem 2.2. Then π : St(k, Rn ) → Gr(k, Rn ) as defined in Eq. (2) is differentiable.
Furthermore, with the charts at hand we obtain that σA is an immersion for any
A ∈ St(k, Rn ).
Proof. First, differentiability with respect to the charts is obvious. We need to show
that ∂(id ◦ σA ◦ ϕ−1A )(ϕA (L)) ∈ L(R
(n−k)×k , Rn×k ) is injective for any L ∈ U . By
A
Theorem 2.2 we have that ϕ−1 A = p A and that σA ◦ pA = (K →
7 A + A ⊥ K), which has
the constant derivative A⊥ . By construction, A⊥ has full rank and the injectivity is
proved.
11
3.1 Proposition. The function π together with the GL(k, R)-actions defined in Eqs.
(13) and (14) is a principal GL(k, R)-bundle.
(i) π is a fiber bundle, i.e. the “local triviality” is satisfied: for every L ∈ Gr(k, Rn )
there is an open neighborhood U of L in Gr(k, Rn ) and a diffeomorphism
ψ : π −1 [U ] → U × GL(k, R),
ψ
π −1 [U ] > U × GL(k, R)
π
∨ proj0
<
U
commutes; here, proj0 denotes the projection onto the first component;
(ii) for every L ∈ Gr(k, Rn ) its fiber π −1 [{L}] is diffeomorphic to GL(k, n), and
ψA (BP ) = (π(BP ), (A> A)−1 A> BP ) = (π(B), (A> A)−1 A> BP ) = ψA (B)P.
4 Riemannian Structure
The Stiefel manifold can be endowed with the structure of a Riemannian manifold by
the Riemannian metric
1
hX, Y iA := tr((A> A)−1 X > Y ) 2 , (15)
where A ∈ St(k, Rn ) and X, Y ∈ TA St(k, Rn ). Note that this is different from the Eu-
clidean Riemannian structure we will also consider later. Since the Grassmann manifold
12
can be equivalently considered as a quotient manifold with respect to the Stiefel man-
ifold modulo GL(k, R) (see Eq. (4)), it can inherit the Riemannian structure from the
Stiefel manifold by turning it into a Riemannian quotient manifold. Since subspaces are
represented by matrices that span the subspaces, the aim of this section is to find rep-
resentations for notions connected with the tangent bundle of the Grassmann manifold
by objects connected with the tangent bundle of the Stiefel manifold. To this end, the
two notions of vertical bundle and horizontal bundle are introduced as follows.
Let A ∈ St(k, Rn ). Since St(k, Rn ) is open in Rn×k one has that TA St(k, Rn ) =
TA Rn×k = Rn×k . Next we decompose the tangent bundle T St(k, Rn ) into two sub-
bundles, the aforementioned vertical and horizontal bundle. First, observe that by
Lemma 1.1 π is surjective and as a consequence of the fiber bundle property π is a sub-
mersion. Hence, by the Submersion Theorem each fiber is a submanifold of dimension
k 2 . The well-defined tangent space to the fiber π −1 [{π(A)}] is a subspace of TA St(k, Rn )
and is called vertical space at A, denoted by VA , i.e.
VA = ker Dπ A ,
since π is constant on fibers. The horizontal space HA at A is, in the case of Riemannian
manifolds considered here, the orthogonal complement in T St(k, Rn ) to VA . This yields
the tangent space to the local cross section through A, i.e.
n o
HA := VA⊥ = Y ∈ T St(k, Rn ); A> Y = 0 ∼ = A⊥ [R(n−k)×k ] ∼= TA SA ,
i.e. all matrices “orthogonal” to A. We denote by V St(k, Rn ) and H St(kn) the vertical
and horizontal bundle associated to St(k, Rn ), respectively.
4.1 Lemma. For any A ∈ St(k, Rn ) one has TA St(k, Rn ) = VA ⊕ HA .
Proof. This is easily verified making use of the full rank of A and A⊥ and their orthog-
onality.
which is meant fiberwise and abbreviates the statement of Lemma 4.1. By Theo-
rem 2.2 we have that Gr(k, Rn ) is a differentiable manifold of dimension k(n − k).
Hence, the tangent space at an arbitrary point L ∈ Gr(k, Rn ) is a vector space of
the same dimension. By Lemma 2.4 σA is an immersion and, hence, the differential
DσA : T UA ⊂ T Gr(k, Rn ) → T SA ⊂ HA ⊂ T St(k, Rn ) is injective. Furthermore, the
tangent spaces have equal dimension so that DσA is a vector space isomorphism between
the tangent space Tπ(A) Gr(k, Rn ) and the horizontal space HA of the Stiefel manifold.
We denote by πGr and πSt the tangent bundle projections of Gr(k, Rn ) and
St(k, Rn ), respectively.
13
4.2 Definition (Horizontal lift). We define
(i) X ◦ π := (A 7→ X(π(A))A ) ∈ X (St(k, Rn )), i.e. the horizontal lift of a vector field
on Gr(k, Rn ) is a vector field on St(k, Rn ).
(ii) Dπ A Y A = Y , or, equivalently, Dπ A restricted to HA is the inverse to ·A . By
definition, that means that Y A is π-related to Y .
(iii) X AP = X A P .
(iv) A> Y AP = 0.
(v) Y f ∼
= ∂Y A (f ◦ π)(A).
Proof. (i) This holds by the definition of the push-forward and by Lemma 2.1(v).
(ii) By Lemma 2.1(iv) we have that
idUA = π ◦ σA .
(iii) This follows from the homogeneity property of σA proved in Lemma 2.1(iv).
(iv) This is a paraphrase of the orthogonality of A and HA .
(v) First observe that f ◦π ∈ F(St(k, Rn )). In submanifolds of Rl derivations (tangent
vectors) and directional derivatives can be identified:
(ii)
∂Y A (f ◦ π)(A) ∼
= Y A (f ◦ π) = Dπ A Y A (f ) = Y f.
Next we define a Riemannian metric on Gr(k, Rn ), which is in fact induced by the
horizontal lift and the Riemannian metric on St(k, Rn ) given in Eq. (15). It is natural
in the following sense: it is the only Riemannian metric that turns π into a Riemannian
submersion, i.e. a submersion whose restriction of the differential to the horizontal bundle
is an isometry; see Eq. (16) in connection with Proposition 4.3(ii).
14
4.4 Proposition (Riemannian metric). For L ∈ Gr(k, Rn ), A ∈ π −1 [{L}] ⊂ St(k, Rn )
and X, Y ∈ TL Gr(k, Rn ) we define
> 1
hX, Y iL = hX, Y iπ(A) := tr((A> A)−1 X A Y A ) 2 = X A , Y A A
. (16)
and is π-related to [X, Y ]Gr (π(A)). These are the characterizing properties of
[X, Y ]Gr (π(A))A and Eq. (17) is proved.
In the following we endow Gr(k, Rn ) with the Riemannian structure induced by the
Riemannian metric given in Eq. (16).
The gradient of a function f ∈ F(Gr(k, Rn ) with respect to the Riemannian metric,
denoted by gradGr f , is the vector field satisfying hgradGr f, Xi = Xf for any X ∈
X (Gr(k, Rn )). Recall that in St(k, Rn ) the Euclidean,
p i.e. induced by Rn×k as in Eq.
>
(1), Riemannian metric is given by hA, BiSt = tr(A B). Thus, the Euclidean gradient
for g ∈ F(St(k, Rn )) with
15
is characterized by
X(f ◦ π) ∼
= ∂X (f ◦ π)(A) = tr(X > gradSt (f ◦ π)(A))
for A ∈ St(k, Rn ), X ∈ TA St(k, Rn ) and f ∈ F(Gr(k, Rn )). Alternatively, we can define
another gradient with respect to the metric defined in Eq. (15) and denoted by gradA f
accordingly. Note that for f ∈ F(Gr(k, Rn )) we have that f ◦ π is constant on fibers and
hence it follows that
hgradA (f ◦ π)(A), XiA = ∂X (f ◦ π)(A) = 0
for all X ∈ VA . Consequently, gradA (f ◦ π)(A) ∈ HA .
4.6 Lemma (Gradient). Let A ∈ St(k, Rn ) and f ∈ F(Gr(k, Rn )). Then
gradGr f (π(A))A = gradA (f ◦ π)(A) = gradSt (f ◦ π)(A)A> A. (19)
Proof. Since gradA (f ◦ π)(A) ∈ HA it suffices to consider horizontal tangent vectors,
that can be uniquely represented as the horizontal lift X A of a tangent vector X ∈
Tπ(A) Gr(k, Rn ). Consider the following equalities:
Now, from Eqs. (21) and (22) we read off the second claimed equality, and from Eqs.
(20) and (23) the first one.
16
Proof. Obviously, both sides of Eq. (24) are elements of HA . Eq. (24) holds if both sides
have the same scalar product with every horizontal tangent vector, which can be repre-
sented by Z A for Z ∈ Tπ(A) Gr(k, Rn ). Since the vertical component of
∇St X(π(A))A , Y ◦ π is orthogonal to the horizontal space and hence there is no con-
tribution to the scalar product, it suffices to prove for each Z ∈ Tπ(A) Gr(k, Rn ) that
D E
∇St X(π(A))A , Y ◦ π , Z A = h∇Gr (X(π(A)), Y ), Ziπ(A) . (26)
A
This follows by expanding both sides in the Koszul formula, where it suffices to consider
terms of the following two types. Let X, Y, Z ∈ X (Gr(k, Rn ) and denote by X ◦ π :=
(A 7→ X ◦ π A ) ∈ X (St(k, Rn )) the horizontal vector field and consider the following as
functions on St(k, Rn ):
X ◦ π Y ◦ π, Z ◦ π = A 7→ X ◦ π A Y ◦ π, Z ◦ π ∼
A
h i
= A 7→ X ◦ π A hY ◦ π, Z ◦ πiπ(·)
∼A
= A 7→ X ◦ π A [hY, Zi· ◦ π]∼A
= A 7→ Dπ A X ◦ π A [hY, Zi]∼π(A)
= A 7→ X(π(A)) [hY, Zi]∼π(A)
= (X hY, Zi) ◦ π,
X ◦ π, Y ◦ π, Z ◦ π = A 7→ X ◦ π A , Y ◦ π, Z ◦ π St (A) A
D E
= A 7→ X ◦ π A , [Y, Z]Gr (π(A))A
A
= A 7→ hX(π(A)), [Y, Z]Gr (π(A))iπ(A)
= (hX, [Y, Z]i) ◦ π.
In summary, corresponding terms in the expansion in the Koszul formula of both sides
of Eq. (26) equal each other and the proof is finished.
In the following we consider regular curves on Gr(k, Rn ). We do not state the domains
explicitly, but assume implicitly that they contain 0.
Let t 7→ C(t) ∈ Gr(k, Rn ) be a regular curve on Gr(k, Rn ). Recall that a vector field
X ∈ XC (Gr(k, Rn )) along C is said to be parallel transported along C if ∇X = 0. A
C
(regular) curve A : t 7→ c(t) ∈ St(k, Rn ) on St(k, Rn ) is called horizontal if Ȧ(t) ∈ HA(t)
for any t ∈ D(A).
Let t 7→ C(t) ∈ Gr(k, Rn ) be a regular curve on Gr(k, Rn ), π(A0 ) = C(0) ∈ Gr(k, Rn ),
A0 ∈ St(k, Rn ). Then there exists a unique horizontal curve A : t 7→ A(t) such that
A(0) = A0 and π(A(t)) = C(t) for all t ∈ D(C). The reason for that is that locally
the image of C is a submanifold in Gr(k, Rn ). For simplicity assume that this holds
17
globally. The preimage of C under π is by the Submersion Theorem a submanifold of
St(k, Rn ), on which we can define a horizontal vector field that is constant on fibers, i.e.
for B ∈ π −1 [C[D(C)]] ⊆ St(k, Rn ) define
X(B) := Ċ(π(B))B .
For each A0 there exists a unique integral curve A through A0 , i.e. a solution to A(0) = A0
and Ȧ(t) = X(A(t)), satisfying π(A(t)) = C(t) for all t ∈ D(C). By definition of X
we have that A is horizontal and the projection property follows from uniqueness. The
curve A is called the horizontal lift of C through A0 .
X A := t 7→ X(t)A(t) ∈ XC (St(k, Rn ))
is the horizontal lift of X along A and X is parallel transported along C if and only if
where
˙ (t) = (s 7→ X(s)
X 0
A A(s) ) (t).
˙ (t) ∈ V
i.e. X A A(t) . Consequently it is of the form
A(t)> X A (t) = 0.
18
Differentiation of the last equation yields
˙ (t) = 0.
Ȧ(t)> X A (t) + A(t)> X A
Ȧ0 (A>
0 A0 )
−1/2
= U ΣV >
be a thin singular value decomposition (SVD), i.e. U ∈ St∗ (k, Rn ), V ∈ St∗ (k, k) and
Σ ∈ Rk×k is diagonal with nonnegative entries; see, for instance, [9, Section 2.5.4].
Then
C(t) = π(A0 (A>
0 A0 )
−1/2
V cos(tΣ) + U sin(tΣ)).
We define exp(Ċ0 ) := C(1).
Proof. Let t ∈ D(C) and A be the unique horizontal lift of C through A0 , so that
Ȧ(t) = Ċ(t)A(t) . Then by (27) we have
−1
Ä(t) + A(t) A(t)> A(t) Ȧ(t)> Ȧ(t) = 0. (29)
Thus, we have
·
\
(A > A) (t) = Ȧ(t)> A(t) + A(t)> Ȧ(t) = 0,
saying that t 7→ A(t)> A(t) is a constant function. Differentiation of Eq. (30) yields
or equivalently
Ȧ(t)> Ȧ(t) = −A(t)> Ä(t). (31)
By plugging Eq. (31) into Eq. (29) we get
−1
Ä(t) − A(t) A(t)> A(t) A(t)> Ä(t) = Ππ(A(t))⊥ Ä(t) = 0,
19
saying that Ä(t) ∈ VA(t) and consequently it is of the form
for some M : t 7→ M (t) ∈ Rk×k . With Eqs. (32) and (31) we obtain that
·
i.e. t 7→ Ȧ(t)> Ȧ(t) is constant, too. Now consider the thin SVD
Ȧ0 (A>
0 A0 )
−1/2
= U ΣV > (33)
Ä(t)(A>
0 A0 )
−1/2
+ A(t)(A>
0 A0 )
−1/2
((A>
0 A0 )
−1/2 >
Ȧ )(Ȧ(A>
0 A0 )
−1/2
) = 0.
Ä(t)(A>
0 A0 )
−1/2
+ A(t)(A>
0 A0 )
−1/2
V Σ2 V > = 0,
Ä(t)(A>
0 A0 )
−1/2
V + A(t)(A>
0 A0 )
−1/2
V Σ2 = 0.
B̈(t) + B(t)Σ2 = 0,
or equivalently
A(t)(A>
0 A0 )
−1/2
V = A0 (A>
0 A0 )
−1/2
V cos(tΣ) + Ȧ0 (A>
0 A0 )
−1/2
V Σ−1 sin(tΣ),
since (A>
0 A0 )
−1/2 V ∈ GL(k, R).
20
Proposition 4.9 shows that the Grassmann manifold with the Riemannian structure
given by the Riemannian metric in (16) is complete, i.e. the geodesics are defined on R.
By [14] the Grassmann manifold is connected and hence, by the Hopf-Rinow-theorem
(see, for instance, [3, Theorem 10.4.16]), a complete metric space with respect to the
Riemannian distance function, also referred to as the geodesic distance. However, the
completeness of Gr(k, Rn ) with respect to the Riemannian distance function can be
obtained alternatively with the notion of principal angles; for an introduction see, for
instance, [9, 12.4.3]. The distance function induced by the Riemannian metric (16) is
the 2-norm of the vector of principal angles denoted by θ. On the other hand, the gap
metric Θ corresponds to the sin of the largest principal angle, i.e. sin |θ|∞ . Now, one
easily verifies the (strong) equivalence of the gap metric and the geodesic distance, i.e.
√
2 k
sin |θ|∞ ≤ |θ|2 ≤ sin |θ|∞ ,
π
from which the completeness of Gr(k, Rn ) with respect to the geodesic distance function
follows. For other definitions of distance functions in terms of principal angles see [4,
4.3].
Acknowledgments
I am grateful to Marcus Köhler and Sascha Trostorff for helpful discussions.
References
[1] P.-A. Absil, R. Mahony, and R. Sepulchre. Riemannian Geometry of Grassmann
Manifolds with a View on Algorithmic Computation. Acta Applicandae Mathemat-
icae, 80(2):199–220, 2004.
21
[7] I. C. Gohberg, P. Lancaster, and L. Rodman. Invariant subspaces of matrices
with applications. Classics in applied mathematics. SIAM Society of Industrial and
Applied Mathematics, 2006.
[8] I. C. Gohberg and A. S. Markus. Two theorems on the opening between subspaces
of a Banach space. Uspekhi Mat. Nauk, 89:135–140, 1959. in Russian.
[9] G.H. Golub and C.F.V. Loan. Matrix computations. Johns Hopkins studies in the
mathematical sciences. Johns Hopkins University Press, 3. edition, 1996.
[10] T. Kato. Perturbation Theory for Linear Operators, volume 132 of Grundlehren der
mathematischen Wissenschaften. Springer, reprint of the 2nd edition, 1995.
Nomenclature
In contrast to the usual differential geometric nomenclature we consequently write ar-
guments of functions in brackets instead of as lower indices, cp. evaluations of vector
fields.
22
ΠL orthogonal projection onto L
S unit sphere in Rn
L⊥ orthogonal complement of L
L⊕M direct sum of L and M
Θ
e spherical gap metric
SA local cross section through A
TA set of matrices which span a subspace without any orthogonal
component to A
σA cross section mapping
A⊥ n × (n − k)-matrix that is orthogonal to A
ϕA chart on UA
pA paramtrization for UA
VA vertical space at A
HA horizontal space at A
V St(k, n) vertical bundle
H St(k, n) horizontal bundle
Tp M tangent space of a manifold M in p ∈ M
TM tangent bundle of a manifold M
D pf push-forward/differential of a differentiable function f in p
πGr , πSt tangent bundle projections of Gr(k, n) and St(k, n), resp.
♦ horizontal lift
Fp (M ) set of functions germs in p ∈ M on a manifold M
F(M ) set of smooth real-valued functions defined on a manifold M
X (M ) set of (smooth) vector fields on a manifold M
X, Y derivations/vector fields
[X, Y ] Lie bracket of X and Y
h·, ·iL Riemannian metric on TL Gr(k, n)
tr trace operator
gradGr , gradSt gradients on Gr(k, n) and St(k, n), resp.
∇(X, Y ) (Riemannian) connection, covariant derivative of Y
in the direction of X
∇St Riemannian connection on St(k, n), directional derivative
∇Gr Riemannian connection on Gr(k, n)
C regular curve on Gr(k, n)
∇X induced covariant derivative for vector fields along C,
C
cp. ∇
dt in [3, Definition 10.1.10]
c (horizontal) regular curve on St(k, n)
D(C) domain of C
23