0% found this document useful (0 votes)
3 views116 pages

Hilbert Space Theory and Its Applications in

The document discusses Hilbert Space Theory and its applications in semi-nonparametric modeling and inference, emphasizing the representation of Borel measurable functions through orthonormal functions. It outlines the foundational concepts of vector spaces, inner products, and projections within Hilbert spaces, and highlights significant models such as the mixed proportional hazard model. The manuscript aims to provide a comprehensive understanding of Hilbert spaces and their relevance in statistical modeling and econometrics.

Uploaded by

hamza.gbada
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views116 pages

Hilbert Space Theory and Its Applications in

The document discusses Hilbert Space Theory and its applications in semi-nonparametric modeling and inference, emphasizing the representation of Borel measurable functions through orthonormal functions. It outlines the foundational concepts of vector spaces, inner products, and projections within Hilbert spaces, and highlights significant models such as the mixed proportional hazard model. The manuscript aims to provide a comprehensive understanding of Hilbert spaces and their relevance in statistical modeling and econometrics.

Uploaded by

hamza.gbada
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

See discussions, stats, and author profiles for this publication at: [Link]

net/publication/267768820

Hilbert Space Theory and Its Applications in

Article · December 2010

CITATION READS
1 4,083

1 author:

Herman J. Bierens
Pennsylvania State University
114 PUBLICATIONS 3,827 CITATIONS

SEE PROFILE

All content following this page was uploaded by Herman J. Bierens on 02 May 2015.

The user has requested enhancement of the downloaded file.


1

Hilbert Space Theory and Its Applications to


Semi-Nonparametric Modeling and Inference

Herman J. Bierens
Pennsylvania State University

Current version: January 25, 2012


2
Chapter 1

Introduction

As is well known, every vector in a Euclidean space can be represented as


a linear combination of orthonormal vectors. Similarly, using Hilbert space
theory, we can represent certain classes of Borel measurable functions1 by
countable infinite linear combinations of orthonormal functions, which allows
us to approximate these functions arbitrarily close by finite linear combina-
tions of these orthonormal functions. This is the basis for semi-nonparametric
(SNP) modeling, where only a part of the model involved is parametrized,
and the non-specified part is an unknown function which is approximated
by a series expansion. See for example Chen (2007) for a recent survey, and
Bickel et al (1998). There is also a substantial literature on estimation of
semi-nonparametric models using nonparametric kernel density and/or re-
gression estimators (see for example Horowitz 1998), but these approaches
are beyond our scope.
Gallant (1981) was the first econometrician to proposed Fourier series
expansions as a way to model unknown functions. Gallant’s approach is
actually nonparametric in that no Euclidean parameters are involved. See
also Eastwood and Gallant (1991) and the references therein. However, the
use of Fourier series expansions to model unknown functions has been pro-
posed earlier in the statistics literature. See for example Kronmal and Tarter
(1968).
Gallant and Nychka (1987) consider SNP estimation of Heckman’s (1979)
sample selection model, where the bivariate error distribution of the latent
1
See for example Bierens (2004, Ch. 2) for the definition of Borel measurability of
functions.

3
4 CHAPTER 1. INTRODUCTION

variable equations is modeled semi-nonparametrically using an Hermite ex-


pansion of the error density.
Another example of a semi-nonparametric model is the mixed propor-
tional hazard (MPH) model proposed by Lancaster (1979). In this model the
hazard function is the product of three factors, the baseline hazard which
depends only on the duration, the systematic hazard which is a function
of the observable covariates, and an unobserved non-negative random vari-
able representing neglected heterogeneity. Elbers and Ridder (1982) have
shown that under some mild conditions and normalizations the MPH model
is nonparametrically identified. Heckman and Singer (1984) propose to es-
timate the distribution function of the unobserved heterogeneity variable by
a discrete distribution. Bierens (2008) and Bierens and Carvalho (2007) use
orthonormal Legendre polynomials to model semi-nonparametrically the un-
observed heterogeneity distribution of interval-censored mixed proportional
hazard models and bivariate mixed proportional hazard models, respectively.
In chapter 2 I will explain what a Hilbert space is, and provide examples
of non-Euclidean Hilbert spaces, in particular Hilbert spaces of Borel measur-
able functions and random variables. In chapter 3 I will discuss projections
on sub-Hilbert spaces and their properties. One of the results involved is the
famous Wold (1938) decomposition theorem, which will be derived first in
general terms and then for covariance stationary time series. Also, the funda-
mental role of the Wold decomposition in time series analysis and empirical
macro-econometrics will be pointed out.
The main focus of this book, however, is on Hilbert spaces of square
integrable Borel measurable real functions and the various orthonormal se-
quences that span these Hilbert spaces, as the basis for semi-nonparametric
modeling and inference. Therefore, following Hamming (1973), in chapter
4 I will review the various ways one can construct orthonormal polynomials
that span a given Hilbert space of functions. In chapter 5 I will show that
any square integrable Borel measurable real function on the unit interval
can be written as a linear combination of the cosine series {cos (kπu)}∞ k=0 ,
u ∈ [0, 1]. This result is related to classical Fourier analysis, which will also
be reviewed. The significance of this result is that it yields closed form se-
ries representations of arbitrary density and distribution functions, as will be
shown in chapter 6, whereas in the approach of Gallant and Nychka (1987),
which is based on Hermite polynomials, and the approach of Bierens (2008)
and Bierens and Carvalho (2007), which is based on Legendre polynomials,
the computation of their density and distribution functions has to be done
5

iteratively. In chapter 7 I will show how to construct compact metric spaces


of density and distribution functions based on the cosine series expansion.
The applications to semi-nonparametric models, based on Bierens. (2011),
will be added to this manuscript in due course.
Throughout this manuscript the set of positive integers will be denoted
by N, and the set of non-negative integers by N0 . Moreover,
√ the well-known
indicator function will be denoted by I(.), and i = −1.
6 CHAPTER 1. INTRODUCTION
Part I
Hilbert spaces

7
Chapter 2

Introduction to Hilbert spaces

In this chapter I will review the concepts of vector spaces, inner products
and Cauchy sequences, and provide examples of Hilbert spaces.

2.1 Vector spaces


The notion of a vector space should be known from linear algebra:

Definition 2.1. Let V be a set endowed with two operations, the operation
”addition”, denoted by ”+”, which maps each pair (x, y) in V ×V into V, and
the operation ”scalar multiplication”, denoted by a dot (.), which maps each
pair (c, x) in R × V [or C × V] into V. Thus, a scalar is a real or complex
number. The set V is called a real [complex ] vector space if the addition and
multiplication operations involved satisfy the following rules, for all x, y and
z in V, and all scalars c, c1 and c2 in R [C]:
(a) x + y = y + x;
(b) x + (y + z) = (x + y) + z;
(c) There is a unique zero vector 0 in V such that x + 0 = x;
(d) For each x there exists a unique vector −x in V such that x + (−x) = 0;1
(e) 1.x = x;
(f ) (c1 c2 ).x = c1 .(c2 .x);
(g) c.(x + y) = c.x + c.y;
(h) (c1 + c2 ).x = c1 .x + c2 .x.

1
Also denoted by x − x = 0.

9
10 CHAPTER 2. INTRODUCTION TO HILBERT SPACES

It is trivial to verify that the Euclidean space Rn is a real vector space.


However, the notion of a vector space is much more general. For example,
let V be the space of all continuous functions on Rn , with pointwise addition
and scalar multiplication defined the same way as for real numbers. Then it
is easy to verify that this space is a real vector space.
Another (but weird) example of a vector space is the space V of positive
real numbers endowed with the ”addition” operation x + y = x.y and the
”scalar multiplication” c.x = xc . In this case the null vector 0 is the number
1, and −x = 1/x.

Definition 2.2. A subspace V0 of a vector space V is a non-empty subset of


V which satisfies the following two requirements:
(a) For any pair x, y in V0 , x + y is in V0 ;
(b) For any x in V0 and any scalar c, c.x is in V0 .

Thus, a subspace V0 of a vector space is closed under linear combinations:


any linear combination of elements in V0 is an element of V0 .
It is not hard to verify that a subspace of a vector space is a vector space
itself, because the rules (a) through (h) in Definition 2.1 are inherited from
the ”host” vector space V. In particular, any subspace contains the null
vector 0, as follows from part (b) of Definition 2.2 with c = 0.

2.2 Inner product and norm


As is well-known, in a Euclidean space Rn the inner product of a pair of
vectors x and y is defined as x0 y, which is a mapping Rn × Rn → R with the
following properties:
(a) x0 y = y 0 x,
(b) (cx)0 y = c(x0 y) for arbitrary c ∈ R,
(c) (x + y)0 z = x0 z + y 0 z,
(d) x0 x > 0 if and only if x 6= 0. √
Moreover, the norm of a vector x ∈ Rn is defined as ||x|| = x0 x. Of course,
in R the inner product is the ordinary product x.y.
Mimicking these four properties, we can define more general inner prod-
ucts with associated norms as follows.
2.2. INNER PRODUCT AND NORM 11

Definition 2.3. An inner product on a real vector space V is a real function


hx, yi: V × V → R such that for all x, y, z in V and all c in R,
(1 ) hx, yi = hy, xi
(2 ) hcx, yi = c hx, yi
(3 ) hx + y, zi = hx, zi + hy, zi
(4 ) hx, xi > 0 if and only if x 6= 0.
An inner product on a complex vector space is defined similarly. The inner
product is then complex-valued, hx, yi: V × V → C. Condition (1 ) then
becomes
(1 *) hx, yi = hy, xi,2
and (2 ) now holds for all complex and real numbers c. Note that also in this
case hx, xi is real valued.3 A vector space endowed with an inner product
is calledpan inner product space. Finally, the norm of x in V is defined as
||x|| = hx, xi

For example, in the vectorR 1 space C[0, 1] of continuous real functions on


[0, 1], the integral hf, gi = 0 f (u) g (u) du is an inner product, with norm
qR
1
kf k = 0
f (u)2 du. Moreover, in the vector space of zero-mean random
variables with finite second moments p the covariance hX, Y i = E[X.Y ] is an
inner product, with norm kXk = E[X 2 ].
As is well-known from linear algebra, for vectors x, y ∈ Rn , |x0 y| ≤
||x||.||y||, which is known as the Cauchy-Schwarz inequality. This inequality
carries over to general inner products:

Theorem 2.1. (Cauchy-Schwarz inequality) |hx, yi| ≤ ||x||.||y||.


p
Given the norm ||x|| = hx, xi, the following properties hold:

||x|| > 0 if x 6= 0; (2.1)


||c.x|| = |c|.||x||; (2.2)
||x + y|| ≤ ||x|| + ||y||. (2.3)

The latter is known as the triangular inequality.


2
The bar denotes the complex conjugate: for z = a + i.b, z = a − i.b.
3
Because hx, xi = hx, xi implies that hx, xi ∈ R.
12 CHAPTER 2. INTRODUCTION TO HILBERT SPACES

The properties (2.1) and (2.2) follow trivially from Definition 2.3. In the
case of a real vector space the triangular inequality (2.3) follows from
||x + y||2 = hx + y, x + yi = hx, xi + 2 hx, yi + hy, yi
= ||x||2 + 2 hx, yi + ||y||2 ≤ ||x||2 + 2 |hx, yi| + ||y||2
≤ ||x||2 + 2||x||.||y|| + ||y||2 = (||x|| + ||y||)2
where the last inequality is due to Theorem 2.1.
In a Euclidean space, a pair x, y of vectors is orthogonal if x0 y = 0, and
orthonormal if also ||x|| = ||y|| = 1. Similarly,

Definition 2.4. Elements x and y in a inner product space with associated


norm are orthogonal if hx, yi = 0, which is also denoted by x ⊥ y, and are
orthonormal if in addition ||x|| = ||y|| = 1.

A norm can also be defined directly:

Definition 2.5. A norm on a vector space V is a mapping ||.||: V → [0, ∞)


such that for all x and y in V and all scalars c the properties (2.1), (2.2) and
(2.3) hold. A vector space endowed with a norm is called a normed space.

2.3 Metric spaces


A norm ||.|| defines a metric d(x, y) = ||x − y|| on V, i.e., a function that
measures the distance between two elements x and y of V, for which (trivially)
the following four properties hold. For all x, y and z in V,
d(x, y) = d(y, x) (2.4)
d(x, y) > 0 if x 6= y; (2.5)
d(x, x) = 0; (2.6)
d(x, z) ≤ d(x, y) + d(y, z). (2.7)
Again, the property (2.7) is known as the triangular inequality.
A metric can also be defined directly:

Definition 2.6. A metric on a space M is a mapping d(., .): M × M →


[0, ∞) satisfying the properties (2.4) through (2.7) for all x, y and z in M.
A space endowed with a metric is called a metric space.
2.4. CONVERGENCE OF CAUCHY SEQUENCES 13

In this definition the space M is not necessarily a vector space: Any space
endowed with a metric is a metric space. For example, let M be the space
of density functions on [0, 1], endowed with the metric
Z 1 ³p p ´2
d(f, g) = f (u) − g (u) du.
0

This space is not a vector space, and it is not possible to define an inner
product on it.

2.4 Convergence of Cauchy sequences


A vector
pspace V endowed with an inner product hx, yi and associated norm
||x|| = hx, yi and metric ||x − y|| is called a pre-Hilbert space. The reason
for the ”pre” is that a fundamental property is still missing, namely that
every Cauchy sequence has a limit in V.

Definition 2.7. A sequence of elements xn of a metric space with metric


d(., .) is called a Cauchy sequence if for every ε > 0 there exists an n0 (ε)
such that for all k, m ≥ n0 (ε), d(xk , xm ) < ε.

For example, in the Euclidean space Rp with finite dimension p every


Cauchy sequence converges to a limit in Rp , and the same applies to the space
Cp of p-dimensional vectors with complex-valued components, endowed with
the inner product

hx, yi = x0 y = (Re (x) − i. Im (x))0 (Re(y) + i. Im(y)) (2.8)


¡ ¢
= Re (x)0 Re(y) + Im (x)0 Im(y)
+i. (Re(x)0 Im(y) − Im(x)0 Re(y))

and associated norm and metric. It is an easy exercise to check that (2.8)
satisfies the conditions in Definition 2.3. Thus,

Theorem 2.2. Every Cauchy sequence in Rp or Cp has a limit in that space.

To demonstrate the role of the Cauchy convergence property, consider


the space C[0, 1] of continuous real functions on [0, 1], i.e., each f ∈ C[0, 1] is
14 CHAPTER 2. INTRODUCTION TO HILBERT SPACES

continuous on (0, 1), and f (0) = limu↓0 f (u) and f (1) = limu↑1 f (u) are finite.
R1
Endow this space with the inner product hf, gi = 0 f (u)g(u)du and associ-
p
ated norm ||f || = hf, f i and metric ||f − g||. Now consider the following
sequence of functions in C[0, 1]:

⎨ 0 for 0 ≤ u < 0.5
n
fn (u) = 2 (u − 0.5) for 0.5 ≤ u < 0.5 + 2−n

1 for 0.5 + 2−n ≤ u ≤ 1,
n = 1, 2, 3, .....
R1
It is an easy
¡ calculus
¢ exercise to verify that ||fk −fm ||2 = 0 (fk (u) − fm (u))2
du < 13 2−k + 2−m , hence fn is a Cauchy sequence in C[0, 1]. Moreover, if
follows from the bounded convergence theorem that limn→∞ ||fn − f || = 0,
where f (u) = I (u > 0.5). However, this limit f (u) is discontinuous in u =
0.5, and thus f ∈ / C[0, 1]. Therefore, the space C[0, 1] is not closed under
convergence.

2.5 Hilbert spaces and sub-Hilbert spaces


2.5.1 Hilbert spaces versus Banach spaces
It is usually quite easy to define an inner product on a vector space, and
the same vector space can often be endowed with different inner products.
For example, for the space of square integrable BorelR measurable functions
1
on [0, 1] we can define an inner product by hf, gi = 0 f (u)g(u)du but also
R1
by hf, gi = 0 uf (u)g(u)du, for example. However, to make such a space
a Hilbert space the inner product must be chosen such that every Cauchy
sequence converges to a limit in that space. The requirement that every
Cauchy sequence in a Hilbert space has a limit in that space makes a Hilbert
space closed under convergence, which generates all kinds of useful properties,
similar to Euclidean spaces.

Definition 2.8. A Hilbert space H is a vector space endowed with an inner


product and associated norm and metric such that every Cauchy sequence in
H has a limit in H. The way the inner product hx, yi is defined, together
with the associated norm and metric, will be called the topology of H.
2.5. HILBERT SPACES AND SUB-HILBERT SPACES 15

Note that the limit of a Cauchy sequence in a Hilbert space is unique. To


see this, suppose that a Cauchy sequence xn ∈ H has two limits: limn→∞ ||xn −
x|| = 0 and limn→∞ ||xn − x∗ || = 0. Then kx∗ − xk = kx∗ − xn + xn − xk ≤
||xn − x|| + ||xn − x∗ || → 0 as n → ∞. Conversely, any convergent sequence in
a Hilbert space is a Cauchy sequence, because limn→∞ ||xn − x|| = 0 implies
that ||xk − xm || = ||(xk − x) + (x − xm )|| ≤ ||xk − x|| + ||xm − x|| → 0 as
min(k, m) → ∞.
In Definition 2.5 the norm ||.|| was defined directly, without reference
to an inner product, giving rise to the definition of a normed space N , for
example. If we endow N with the metric ||x − y|| and require that every
Cauchy sequence in N has a limit in N then the space N becomes a Banach
space. The difference between a Hilbert space and a Banach space is the
source of the norm: In an Hilbert space the norm is defined on the basis of
an inner product whereas in the case of a Banach space the norm is defined
directly. Consequently, in a Banach space the notion of inner product is
nonexisting, and so is the notion of orthogonality.

2.5.2 Linear manifolds and sub-Hilbert spaces


Because a Hilbert space is a vector space, we can define a subspace of a
Hilbert space in the same way as for vector spaces (see Definition 2.2), and
endow it with the same inner product, norm and metric as Hilbert space
involved. Such a subspace is called a linear manifold:

Definition 2.9. A linear manifold M of a Hilbert space H is a subspace of


H endowed with the topology of H.

However, a linear manifold M is not necessarily a Hilbert space itself. In


general there is no guarantee that every Cauchy sequence in M takes a limit
in M. If so the linear manifold M needs to be extended by augmenting it
with the limits of all Cauchy sequence in M. The resulting extended linear
manifold coincides with the closure M of M. Recall that M is a subset
of the metric space H, and that a point of closure of M is an element x
such that for each ε > 0 there exists a z ∈ M and a y ∈ H\M such that
||x − z|| < ε and ||x − y|| < ε. The set of all points of closure of M is called
the border of M, denoted by ∂M, and the closure of M, denoted by M, is
the union of M and its border: M = M ∪ ∂M.
16 CHAPTER 2. INTRODUCTION TO HILBERT SPACES

Theorem 2.3. The closure M of a linear manifold M is a Hilbert space.

In other words, M is a sub-Hilbert space.

2.5.3 Hilbert spaces spanned by a sequence


Let H be a Hilbert space and let {xk }∞ k=1 be a sequence of elements of H.
Let Mm be the linear manifold spanned by x1 , ..., xm , i.e., Mm consists of
all linear combinations of x1 , ..., xm . Then it follows similar to the proof of
Theorem 2.3 that

Lemma 2.1. Mm is a Hilbert space.

Definition 2.10. The space M∞ = ∪∞ n=1 Mn is called the space spanned by


{xj }∞
j=1 , and is also denoted by span({x ∞
j }j=1 ).

It follows similar to the proof of Theorem 2.3 that

Lemma 2.2. M∞ is a Hilbert space.

Remark. In the sequel a sub-Hilbert space will be referred to as a ”sub-


space”.

Definition 2.11. A sequence {xk }∞


k=1 in a Hilbert space H is called complete

if H = span({xj }j=1 ).

2.6 Examples of non-Euclidean Hilbert spaces


2.6.1 A Hilbert space of random variables
Consider the space R of random variables defined on a common probability
space {Ω, F, P } with finite second moments, endowed
p with the inner
p product
hX, Y i = E [X.Y ] and associated norm ||X|| = hX, Xi = E[X 2 ] and
metric ||X − Y ||. Then

Theorem 2.4. The space R is a Hilbert space.


2.7. APPENDIX: PROOFS 17

2.6.2 Hilbert spaces of functions


Let w(x) be a probability density on R and let L2 (w) be the space of Borel
measurable real functions f on R satisfying
Z ∞
f (x)2 w(x)dx < ∞
−∞

where the integral is the Lebesgue integral, endowed with the inner product
Z ∞
hf, gi = f (x)g(x)w(x)dx
−∞

p qR

and associated norm ||f || = hf, f i = −∞
f (x)2 w(x)dx and metric
||f − g||. Then for f, g ∈ L2 (w) , hf, gi = E [f (X)g(X)] , where X is a
random drawing from the distribution with density w(x), hence it follows
from Theorem 2.4 that L2 (w) is a Hilbert space.

2.7 Appendix: Proofs


2.7.1 Theorem 2.1
Let the vector space involved be complex. It follows from the properties
(1)-(4) in Definition 2.3 that for any complex valued λ,

0 ≤ hx + λy, x + λyi
= hx, xi + hλy, xi + hx, λyi + hλy, λyi
= ||x||2 + λ hy, xi + hλy, xi + λ hy, λyi
= ||x||2 + λhx, yi + λ hy, xi + λhλy, yi
= ||x||2 + λhx, yi + λ.hy, xi + λ.λ hy, yi
= ||x||2 + λhx, yi + λ. hx, yi + λ.λ hy, yi
= ||x||2 + λhx, yi + λ. hx, yi + λ.λ||y||2

Next, note that

λhx, yi + λ. hx, yi = (Re (λ) + i. Im (λ)) (Re (hx, yi) − i. Im (hx, yi))
+ (Re (λ) − i. Im (λ)) (Re (hx, yi) + i. Im (hx, yi))
= 2 (Re (λ) Re (hx, yi) + Im (λ) Im (hx, yi))
18 CHAPTER 2. INTRODUCTION TO HILBERT SPACES

and

λ.λ = (Re (λ) + i. Im (λ)) (Re (λ) − i. Im (λ))


= (Re (λ))2 + (Im (λ))2

Hence

0 ≤ ||x||2 + 2 (Re (λ) Re (hx, yi) + Im (λ) Im (hx, yi)) (2.9)


¡ ¢
+ (Re (λ))2 + (Im (λ))2 .||y||2

The latter is minimal for


Re (hx, yi) Im (hx, yi)
Re (λ) = − 2
, Im (λ) = − .
||y|| ||y||2

Substituting these solutions in (2.9) yields

1 ¡ 2 2¢ |hx, yi|2
0 ≤ ||x||2 − (Re (hx, yi)) + (Im (hx, yi)) = ||x||2

||y||2 ||y||2

and thus |hx, yi| ≤ ||x||.||y||.

2.7.2 Theorem 2.2


Let xn be a Cauchy sequence in R, and denote x = lim supn→∞ xn . Let us
first show that x < ∞, as follows. By the definition of ”limsup” there exists
a subsequence nk such that x = limk→∞ xnk . Note that this xnk is also a
Cauchy sequence, hence for arbitrary ε > 0 there exists an index k0 such
that |xnk − xnm | < ε for all k, m ≥ k0 . Keeping m ≥ k0 fixed and letting
k → ∞ it follows that |x − xnm | < ε, hence x < ∞. By a similar argument
it follows that x = lim inf n→∞ xn > −∞. Thus, we can find ¯ an index ¯ k0
and subsequences n and n such that for all k, m ≥ k , ¯ x − x ¯
n1,m < ε,
¯ ¯ 1,k¯ 2,m ¯ 0
¯x − xn2,m ¯ < ε and ¯xn1,m − xn2,m ¯ < ε, hence |x − x| < 3ε. Since ε was
arbitrary, it follows now that x = x = x, which implies that limn→∞ xn =
x ∈ R. By applying this argument to the real and imaginary parts of a
complex valued Cauchy sequence xn it follows that limn→∞ xn = x ∈ C.
Moreover, applying this argument to each component of a (complex) vector
valued Cauchy sequence the results for the cases Rp and Cp follow.
2.7. APPENDIX: PROOFS 19

2.7.3 Theorem 2.3


Let xn be a Cauchy sequence in M ⊂ H. Then xn has a limit x ∈ H, i.e.,
limn→∞ ||xn − x|| = 0. Suppose that x ∈ / M. Since M is closed there exists
an ε > 0 such that the set N (x, ε) = {x ∈ H : ||x − x|| < ε} is completely
outside M: N (x, ε) ∩ M = ∅. But limn→∞ ||xn − x|| = 0 implies that there
exists an n(ε) such that xn ∈ N (x, ε) for all n > n(ε), hence xn ∈
/ M for all
n > n(ε), which contradicts xn ∈ M for all n.

2.7.4 Lemma 2.1


Without loss of generality we may assume that the m × m matrix Xm with
elements hxi , xj i is nonsingular, as otherwise we canPre-arrange the xj ’s such
that Mm = Mr with r = rank(Xm ). Let yn,m = m j=1 cj,n xj be a Cauchy
sequence in Mm . Then
°m °2
°X °
° °
||yn1 ,m − yn2 ,m ||2 = ° (cj,n1 − cj,n2 )xj °
° °
j=1
XX
m m
= (ci,n1 − ci,n2 )(cj,n1 − cj,n2 )hxi , xj i → 0
i=1 j=1

as min(n1 , n2 ) → ∞. This is only possible if for j = 1, 2, ..., m, limmin(n1 ,n2 )→∞


|cj,n1 − cj,n2 | = 0. Thus, the cj,n ’s areP
Cauchy sequences in R, and therefore
converge to limits cj . Denoting ym = m j=1 cj xj , which is an element of Mm ,
it follows now easily that limn→∞ ||yn,m − ym || = 0. Thus, every Cauchy
sequence in Mm converges to a limit in Mm .

2.7.5 Theorem 2.4


Let Xn be a Cauchy sequence in R. Then
£ ¤
||Xn − Xm ||2 = E (Xn − Xm )2 → 0

as min(n, m) → ∞, so that by Chebyshev’s inequality,

plim |Xn − Xm | = 0.
min(n,m)→∞
20 CHAPTER 2. INTRODUCTION TO HILBERT SPACES

As is well-known, convergence in probability is equivalent to almost sure


(a.s.) convergence along a further subsequence of an arbitrary subsequence4 ,
hence there exists a subsequence nk such that for min(k, m) → ∞,
a.s.
|Xnk − Xnm | → 0.

In its turn this result is equivalent to the statement that there exists a set
N ∈ F with P (N) = 0, called a null set, such that for all ω ∈ Ω\N,

lim |Xnk (ω) − Xnm (ω)| = 0


min(k,m)→∞

Now Xnk (ω) is a Cauchy sequence in R and thus converges to a limit X(ω) in
R, which is measurable F,5 so that X is a random variable defined on
{Ω, F, P }. Hence, for fixed m and k → ∞
a.s.
(Xnk − Xm )2 → (X − Xm )2 . (2.10)

Finally, it follows from (2.10), Fatou’s lemma6 and the Cauchy property that
£ ¤ h i
||X − Xm ||2 = E (X − Xm )2 = E lim (Xnk − Xm )2
k→∞
£ ¤
≤ lim inf E (Xnk − Xm )2 → 0
k→∞

for m → ∞.

4
See for example Bierens (2004, Theorem 6.B.3, p. 168).
5
The latter follows from the well-known property that the limsup and liminf of a se-
quence of random variables are random variables themselves. See for example Bierens
(2004, Theorem 2.13, p. 47).
6
Fatou’s lemma states: For a sequence Xn of non-negative random variables,
E[lim inf n→∞ Xn ] ≤ lim inf n→∞ E[Xn ]. See for example Bierens (2004, Lemma 7.A.1,
p. 201).
Chapter 3
Projections

3.1 The projection theorem


As is well-known from linear algebra and econometrics, the projection of a
vector y ∈ Rn on the n
Pk subspace spanned by vectors x1 , ..., xk in R is a linear
combination yb = j=1 βj xj such that ||y − yb|| is minimal. This is a linear
regression problem: Minimize

||y − yb||2 = y 0 y − 2y 0 Xβ + β 0 X 0 Xβ

to β = (β1 , ..., βk )0 , where X = (x1 , ..., xk ) . If k ≤ n and the vectors x1 , ..., xk


are linear independent then the solution is β = (X 0 X)−1 X 0 y, hence yb =
Xβ = X (X 0 X)−1 X 0 y.
If x1 , ..., xk are not linear independent then rank(X) = m < k. In that
case we can rearrange x1 , ..., xk such that the matrix X1 = (x1 , ..., xm ) has
rank m and X2 = (xm+1 , ..., xk ) = X1 C for some (k − m) × (k − m) matrix
C. Partition β accordingly as β = (b01 , b02 )0 . Then

||y − yb||2 = y 0 y − 2y 0 X1 (b1 − Cb2 ) + (b1 − Cb2 )0 X10 X1 (b1 − Cb2 )

which is minimal for (b1 − Cb2 ) = (X10 X1 )−1 X10 y, hence


−1
yb = X1 b1 + X2 b2 = X1 (b1 − Cb2 ) = X1 (X10 X1 ) X10 y,

which is unique. The latter follows from Theorem 3.1 below.


The notion of a projection for Hilbert spaces is similar:

21
22 CHAPTER 3. PROJECTIONS

Definition 3.1. The projection yb of an element y of a Hilbert space H on a


subspace S is an element yb ∈ S such that ||y − yb|| = inf z∈S ||y − z||.

However, we still have to show that yb ∈ S is possible and unique. This


follows from the fundamental projection theorem:

Theorem 3.1. (Projection theorem) If S is a subspace of a Hilbert space


H and y an element of H then there exists a unique element yb ∈ S such that
||y − yb|| = inf z∈S ||y − z||. Moreover the residual u = y − yb is orthogonal to
any z ∈ S: hu, zi = 0.

3.2 Projections in terms of angles


As is well known, the angle ϕ(x, y) between two vectors x and y in a Euclidean
space is defined by the cosine formula

||x||2 + ||y||2 − ||x − y||2 x0 y


cos(ϕ(x, y)) = = ,
2||x||.||y|| ||x||.||y||

due to the Law of Cosines.1 Clearly, this formula carries over to elements x
and y of a Hilbert space
√ H, simply by replacing p
the Euclidean inner product
0 0
x y and norm ||x|| = x x by hx, yi and ||x|| = hx, xi, respectively. Thus,
the angle ϕ(x, y) between two elements x and y of a Hilbert space is defined
by the cosine formula

hx, yi
cos(ϕ(x, y)) = . (3.1)
||x||.||y||

Let S, y and yb be as before, and let x be any element of S. Then it follows


from the cosine formula (3.1) and the orthogonality condition hx, y − ybi =
0 that
hx, yi hx, ybi ||b
y ||
cos(ϕ(x, y)) = = = cos(ϕ(x, yb))
||x||.||y|| ||x||.||y|| ||y||
1
Consider a triangle ABC, let ϕ be the angle between the legs C → A and C → B,
and denote the lengths of the legs opposite to the points A, B and C by α, β, and γ,
respectively. Then γ 2 = α2 + β 2 − 2αβ cos(ϕ).
3.3. PROJECTIONS ON SUBSPACES SPANNED BY A SEQUENCE 23

which is maximal if cos(ϕ(x, yb)) = 1. The latter is true if x = c.b


y for some
constant c > 0. Consequently,
||b
y ||
cos(ϕ(y, yb)) = max cos(ϕ(x, y)) = . (3.2)
x∈S ||y||

3.3 Projections on subspaces spanned by a


sequence
Let H be a Hilbert space and let {xk }∞ k=1 be a sequence of elements of H.
Let Mn be the linear manifold spanned by x1 , ..., xn : Mn = span({xk }nk=1 ) .
As we have seen from Lemma 2.1, Mn is a Hilbert space.
Consider Pthe projection ybn of an element y ∈ H on Mn . Then ybn takes the
n
form ybn = k=1 θn,k xk , where the θn,k ’s are the solutions of the minimization
problem
° °2
° Xn °
° °
min °y − θk xk °
θ1 ,θ2 ,...,θn ° °
k=1
à !
X
n X
n X
n
= min ||y||2 − 2 θk hxk , yi + θk θm hxk , xm i
θ1 ,θ2 ,...,θn
k=1 k=1 m=1

Similar to linear regression, the first-order conditions involved are the normal
equations
Xn
hxk , xm i θn,m = hxk , yi , k = 1, 2, ..., n,
m=1

which can be written in matrix-vector form as Σn,xx θn = Σn,xy , for example.


To solve this system uniquely as θn = Σ−1
n,xx Σn,xy we need to impose a similar
condition as linear independence in Euclidean spaces, namely regularity:

Definition 3.2. Let {xk }∞k=1 be a sequence


¡ of elements
¢ of a Hilbert space H.

Denote the projection of xk on span {xj }j=k+1 by x bk , and let uk = xk − xbk .

The sequence {xk }k=1 is said to be regular if ||uk || > 0 for all k ≥ 1.

Exercise: Given a regular sequence {xk }∞k=1 , prove that for n = 1, 2, 3, .....
the n × n matrices Σn,xx with elements hxi , xj i are nonsingular.
24 CHAPTER 3. PROJECTIONS

Lemma 3.1. For z ∈ span({xk }∞ k=1 ) let zbn be the projection of z on


n
span({xk }k=1 ). Then limn→∞ ||z − zbn || = 0.

More generally we have:

Theorem 3.2. For z ∈ H, let zb be the projection of z on span({xk }∞ k=1 ) and


n
let zbn be the projection of z on span({xk }k=1 ). Then limn→∞ ||b
z − zbn || = 0.

Although each projection zbn is a linear combination of x1 , ..., xn , in gen-


eral the result of Theorem
P∞ 3.2 does not imply that there exists a sequence

{θj }j=1 such that zb = j=1 θj xj . As an example of such a case, consider the
Hilbert space R0 of zero-mean random variables with finite second moments,
endowed with the inner product hX, Y i = E[X.Y ] and associated norm and
metric. Let
Xt = Vt − Vt−1 ,
where Vt is distributed i.i.d. N(0, 1). This is clearly a zero-mean covariance
stationary process, with covariance function γ(0) = 2, γ(1) = −1, γ(m) =
0 for m ≥ 2. Hence Xt ∈ R0 for all t.
For given t, let Mt−1 ∞ t−1
−∞ = span({Xt−m }m=1 ) , Mt−n = span(Xt−1 , ...., Xt−n ) .
The projection X bt,n of Xt on Mt−1
t−n takes the form

X
n
bt,n =
X θn,j Xt−j
j=1

where the coefficients θn,j are the solutions of the normal equations
X
n
γ(m) = γ(|k − m|)θn,k , m = 1, ..., n.
k=1

hence for n ≥ 3,
−1 = 2.θn,1 − θn,2
0 = −θn,1 + 2θn,2 − θn,3
0 = −θn,2 + 2θn,3 − θn,4
..
.
0 = −θn,n−2 + 2θn,n−1 − θn,n
0 = −θn,n−1 + 2θn,n
3.3. PROJECTIONS ON SUBSPACES SPANNED BY A SEQUENCE 25

The solutions of these normal equations are


j
θn,j = − 1, j = 1, ...., n,
n+1
hence n µ ¶
X j
b
Xt,n = − 1 Xt−j (3.3)
j=1
n+1
bt be the projection of Xt on Mt−1
Next, let X , and suppose that there
∞ b P∞ −∞
exists a sequence {θj }j=1 such that Xt = j=1 θj Xt−j . Note that the latter
is merely a short-hand notation for
° °2 ⎡Ã !2 ⎤
° X n ° X
n
°b ° bt −
lim °X t− θj Xt−j ° = lim E ⎣ X θj Xt−j ⎦ = 0 (3.4)
n→∞ ° ° n→∞
j=1 j=1

If so, it follows from Theorem 3.2 and (3.3) that


⎡Ã !2 ⎤
Xn Xn µ ¶
j
0 = lim E ⎣ θj Xt−j − − 1 Xt−j ⎦ (3.5)
n→∞
j=1 j=1
n + 1
⎡Ã !2 ⎤
Xn µ ¶
j
= lim E ⎣ − 1 − θj Xt−j ⎦
n→∞
j=1
n + 1

But
Xn µ ¶ X n µ ¶
j j
− 1 − θj Xt−j = − 1 − θj (Vt−j − Vt−j−1 )
j=1
n + 1 j=1
n + 1
µ ¶ Xµ
n−1 ¶
n 1
=− + θ1 Vt−1 − θj+1 − θj − Vt−j−1
n+1 j=1
n+1
µ ¶
1
+ + θn Vt−n−1
n+1
hence
⎡Ã !2 ⎤ µ
X n µ ¶ ¶2
j n
E⎣ − 1 − θj Xt−j ⎦ = + θ1
j=1
n + 1 n + 1
n−1 µ
X ¶2 µ ¶2
1 1
+ θj+1 − θj − + + θn (3.6)
j=1
n+1 n+1
26 CHAPTER 3. PROJECTIONS

This equality implies that for arbitrary integers m ≥ 1,


⎡Ã !2 ⎤
X n µ ¶
j
lim inf E ⎣ − 1 − θj Xt−j ⎦
n→∞
j=1
n + 1
µ ¶2 µ ¶2
n 1
≥ lim inf + θ1 + lim inf θm+1 − θm −
n→∞ n + 1 n→∞ n+1
2 2
= (θ1 + 1) + (θm+1 − θm ) .

Therefore, a necessary condition for (3.5) is that θm = −1 for m = 1, 2, 3, .....


But then it follows from (3.6) that
⎡Ã !2 ⎤
Xn µ ¶ µ ¶2
j 1
lim E ⎣ − 1 − θj Xt−j ⎦ = lim −1 =1
n→∞
j=1
n + 1 n→∞ n + 1

which contradicts (3.5). Thus, in this case there does not exist a sequence
{θj }∞
j=1 such that (3.4) holds. ¡ ¢
The problem that for the projection zb on span
P {xj }∞j=1 there does not
always exist a sequence {θj }∞ b= ∞
j=1 such that z j=1 θj xj only occurs if the
sequence {xj }∞j=1 is not orthogonal:

Theorem 3.3. If a sequence {xj }∞ j=1 in a Hilbert space H is orthonormal,


i.e.,
hxi , xj i = I(i = j), (3.7)
P ∞
then the projection zb of z ∈ H on span({xj }∞ b = j=1 θj xj
j=1 ) takes the form z
Pn P
(in the sense that limn→∞ ||bz − j=1 θj xj ||), where θj = hz, xj i with ∞ 2
j=1 θj <
∞.

3.4 The Wold decomposition


Let S1 , ..., Sn be subspaces of a Hilbert space H. Then similar to Definition
2.10,

P Span(S1 , ..., Sn ) is the closure of the space of all linear


Definition 3.3.
combinations nj=1 cj xj , where xj ∈ Sj .
3.4. THE WOLD DECOMPOSITION 27

We also need the definition of orthogonal complement:

Definition 3.4. The orthogonal complement of a subspace S of a Hilbert


space H, denoted by S ⊥ , is the subset of H such that for each x ∈ S and
y ∈ S ⊥ , hx, yi = 0.

Lemma 3.2. Orthogonal complements are subspaces.

We can now formulate the following general version of the Wold decom-
position:


Theorem 3.4. Given a regular sequence {xk }P k=1 in a Hilbert space, every

x ∈ S = span({xk }k=1 )Pcan be written as x = ∞ k=1 αk ek + w, in the sense
that limn→∞ kx − w − nk=1 P αk ek k = 0, where {ek }∞
k=1 is an orthonormal
∞ 2
sequence in S, αk = hx, ek i , k=1 αk < ∞, and

w ∈ S∞ ∩ U∞ , (3.8)
with S∞ = ∩∞ ∞ ⊥
n=1 span({xk }k=n ) and U∞ the orthogonal complement of U∞ =

span({ek }k=1 ). Note that (3.8) implies that w is orthogonal to all the ek ’s:
hek , wi = 0 for k = 1, 2, 3, ....

In the case of the Hilbert space R0 of zero-mean random variables with


finite second moments, with inner product hX, Y i = E[X.Y ] and associated
norm and metric, the results of Theorem 3.4 translate as follows:

Theorem 3.5. (Wold decomposition theorem) Let Xt be a regular univariate


zero-mean covariance stationary time series process. Then Xt can be written
as
X∞
Xt = αj Ut−j + Wt a.s., (3.9)
j=0

where Ut is a zero-mean uncorrelated process with variance 1,


X

αj = E[Xt Ut−j ], αj2 < ∞, (3.10)
j=0

and Wt is a zero-mean covariance stationary process satisfying


Wt ∈ Ut⊥ ∩ S−∞ , (3.11)
28 CHAPTER 3. PROJECTIONS

where S−∞ = ∩n span({Xn−k }∞ ⊥


k=1 ) and Ut is the orthogonal complement of
Ut = span({Ut−k }∞
k=0 ) . The result (3.11) implies that

Wt ∈ span ({Wt−m }∞
m=1 ) , (3.12)

which in its turn implies that Wt is perfectly predictable from the past values
Wt−1 , Wt−2 , Wt−3 , ........ Moreover, (3.11) implies that

E[Wt Ut−m ] = 0 (3.13)

for all leads and lags m.

The condition var(Ut ) = 1 is not essential as long as Xt is regular. With-


out loss of generality we may then replace Ut with U et = σUt , σ > 0, and αk
with αe k /σ, where σ can be pinned down by normalizing α e 0 = 1.
The Wold decomposition carries over to k-variate covariance stationary
processes Xt , as follows. Consider the Hilbert space Rk of zero mean random
vectors in Rk with finite second moment matrices, endowed with the inner
product hX, Y i = E [X 0 Y ] and associated norm and metric. Let X bt be the

projection of Xt on span({Xt−j }j=1 ), with residual vector Vt = Xt − X bt , and
0
let Σ = E [Vt Vt ] . In this case we need to extend the notion of regularity
by requiring that Σ is positive definite rather than only ||Vt ||2 = E [Vt0 Vt ] >
0, so that we can define Ut = Σ−1/2P Vt . Then the projection X et of Xt on
£ ¤
n e n
span({Ut−j }j=0 ) takes the form Xt = j=1 Aj Ut−j , where Aj = E Xt Ut−j 0
.
It follows now straightforwardly from the proofs of Theorems 3.4 and 3.5
that
X∞
Xt = Aj Ut−j + Wt a.s.,
j=1

where the process Ut is uncorrelated with zero expectation vector and vari-
ance matrix Ik , and Wt ∈ Ut⊥ ∩ S−∞ , with Ut⊥ and S−∞ defined in Theorem
3.5.
It should be stressed that the deterministic process Wt is not necessarily
nonrandom. For example let Wt = a. cos(λt) + b. sin(λt), where a and b are
independent random drawings from the standard normal distribution and λ ∈
(−π/2, π/2) is a constant. Then E [Wt ] = 0 and E [Wt Wt−m ] = cos (λm) ,
hence Wt is a zero-mean covariance stationary process. If we observe Wt−1 ,
Wt−2 and Wt−3 then we can solve a, b and λ, hence Wt is then determined
for all t.
3.4. THE WOLD DECOMPOSITION 29

The question now arises under which conditions the deterministic process
Wt is identical to zero. Since Wt ∈ ∩n span({Xn−j }∞
j=0 ), it follows that Wt is
measurable with respect to the remote σ-algebra of the process Xt :

Definition 3.5. Let Ft = σ({Xt−j }∞ j=0 ) be the σ-algebra generated by


{Xt−j }∞
j=0 . The σ-algebra F−∞ = ∩ F
t t is called the remote σ-algebra of the
process Xt .

If the process Xt is independent then it follows from Kolmogorov’s zero-


one law2 that the sets in F−∞ have either probability one or zero, so that the
information in F−∞ is non-informative. In other words, the memory of the
remote past of Xt has vanished. However, this result carries over to certain
dependent processes, for example α-mixing processes.3 This gives rise to the
notion of vanishing memory:

Definition 3.6. A time series process is said to have a vanishing memory


if the sets in its remote σ-algebra F−∞ have either probability one or zero,
i.e., A ∈ F−∞ implies P [A] = 1 or P [A] = 0.

In that case E[Wt |F−∞ ] = E[Wt ] a.s.4 However, since Wt is measurable


F−∞ , we also have E[Wt |F−∞ ] = Wt a.s. Thus, Wt = E[Wt ] = 0 a.s., where
the second equality follows from the condition that E[Xt ] = 0. Consequently,

Theorem 3.6. If the zero-mean covariance stationary process Xt has a


vanishing memory then the deterministic term Wt in its Wold decomposition
is zero with probability 1.

The Wold decomposition theorem in the form of Theorem 3.6 is the basis
for time series analysis. In particular, for a univariate covariance stationary
process Xt with a vanishing memory and expectation E[Xt ] = µ the Wold
decomposition can be written as
Xt = µ + α (L) Ut
P
where L is the lag operator and α (L) = 1 + ∞ k
k=1 αk L . The function α (L)
can be approximated arbitrarily close by a ratio of two lag polynomials,
2
See for example Bierens (2004, Theorem 7.5, p.185).
3
See for example Bierens (2004, Theorem 7.6, p.186).
4
See for example Bierens (2004, Exercise 3 in Section 7.6).
30 CHAPTER 3. PROJECTIONS
P P
ψq (L) = 1 + qk=1 θk Lk and ϕp (L) = 1 − pk=1 γk Lk , of orders q and p,
5
respectively, where at least ϕp (L) is invertible with inverse ϕ−1p (L). In
particular, for arbitrary ε > 0 there exist lag polynomials ψq (L) and ϕp (L)
such that h¡¡ ¢ ¢2 i
E α (L) − ϕ−1
p (L) ψ q (L) Ut < ε.
This gives rise to the well-known ARMA(p, q) models, for which it is assumed
that α (L) is exactly of the form α (L) = ϕ−1
p (L) ψq (L) , so that ϕp (L) Xt =
γ + ψq (L) Ut with γ0 = ϕp (1) µ. Thus,
p q
X X
Xt = γ0 + γk Xt−k + Ut + θm Ut−m .
k=1 m=1

Moreover, if also ψq (L) is invertible then Xt has the representation

ψq−1 (L) ϕp (L) Xt = β0 + Ut ,

The lag function ψq−1 (L) ϕp (L) can be writ-


where β0 = µ.ϕp (1) /ψq (1) . P
ten as ψq−1 (L) ϕp (L) = 1 − ∞ k
k=1 βk L , so that then Xt has the AR(∞)
representation
X∞
Xt = β0 + βk Xt−k + Ut .
k=1

This representation plays a key role in forecasting.


An important econometric application of the multivariate version of the
Wold decomposition is Sims’ (1980) innovation response analysis. Sims’
(1980) landmark paper has changed the way empirical macroeconomics is
conducted nowadays. His idea is the following. Let Xt ∈ Rk be a covariance
stationary process of economic variables generated by a stationary VAR(p)
process:
p
X
Xt = b0 + Bk Xt−k + Ut
k=1

Assume that the error vectors Ut are i.i.d. Nk [0, Σ] , where Σ is nonsingular.
Stationarity of this process is equivalent
P to the requirement that the matrix-
valued lag polynomial B(L) = Ik − pk=1 Bk Lk is invertible.6 The latter
5
I.e., ϕp (z) = 0 for some z ∈ C implies |z| > 1.
6
Which in its turn is equivalent to the condition that the roots of the polynomial
det(B(z)) are all outside the complex unit circle: det(B(z)) = 0 implies |z| > 1.
3.5. PROJECTIONS ON A RANDOM SUBSPACE 31

condition also guarantees that Xt has a vanishing memory. It follows then


from the Wold decomposition that Xt can be decomposed as

X

Xt = µ + Am Ut−m ,
m=0

where µ = E[Xt ] and A0 = Ik . The parameters Σ, µ, and Am can be


estimated by estimating the VAR(p) for Xt by ordinary least squares, and
then inverting the VAR(p) lag polynomial.
The variance matrix Σ of the Ut ’s can be written as Σ = ∆.∆0 , where ∆ is
a k × k lower-triangular matrix, so that Ut can be written as Ut = ∆et , where
now et ∼ Nk [0, Ik ] . Sims proposes to interpret the components of et as the
unpredictable parts of policy interventions in the corresponding components
of Xt . To trace the effect of these policy innovations on the future path
of Xt , project Xt+m for m ≥ 0 on component ei,t of et . These projections
take the form Am δi ei,t , where δi is column i of ∆, and may be interpreted as
the response of Xt+m to the innovation ei,t . Since the scale of ei,t does not
matter, the responses of Xt+m for m = 0, 1, 2, .... to a unit shock in ei,t are
now Am δi = E [Xt+m |ei,t = 1] − E [Xt+m ] , which are usually presented in the
form of graphs.
For more on the Wold decomposition and its time series applications, see
for example Anderson (1994).

3.5 Projections on a random subspace


Because a Hilbert space H is a metric space, we can define open sets in H in
the usual way. Therefore, similar to the Euclidean Borel field, we can define
the Borel field BH of subsets of H as the smallest σ-algebra containing the
collection of all open sets in H, and call its elements Borel sets. Moreover,
given a probability space {Ω, F, P }, where F is a σ-algebra of subsets of
the sample space Ω and P is a probability measure on F, a random element
X ∈ H can now be defined as a mapping X(.) : Ω → H such that for all
Borel sets B ∈ BH , {ω ∈ Ω : X(ω) ∈ B} ∈ F. In particular, if Xn is a
sequence of random elements of H defined on a common probability space,
the notion of convergence in probability of Xn to a non-random element x
of H, denoted by plimn→∞ ||Xn − x||, can be defined in the same way as for
32 CHAPTER 3. PROJECTIONS

sequences of random vectors in a Euclidean space, i.e., for an arbitrary ε > 0,

def.
Pr [||Xn − x|| < ε] = P ({ω ∈ Ω : X(ω) ∈ {z ∈ H : ||z − x|| < ε}})
→ 1 as n → ∞,

because {z ∈ H : ||z − x|| < ε} is an open set and therefore a Borel set.
Now the question I will address is the following. Let YN and X1,N , ..., Xn,N
be random elements of H depending on a sample of size N, where n = nN is a
subsequence of N. Let Ybn,N be the projection of YN on span(X1,N , ..., Xn,N ),
and let Un,N = YN − YbN be the residual. Under what conditions do Ybn,N
and Un,N converge in probability? The answer to this question is crucial for
proving asymptotic normality of semi-nonparametric sieve estimators.
The answer is given in the following theorem.

Theorem 3.7. Let YN and X1,N , X2,N , ...., Xn,N be random elements of a
Hilbert space H on the basis on a sample of size N, where n is a subsequence
of N. Let Ybn,N be the projection of YN on span({Xm,N }nm=1 ) , with residual
Un,N = YN − Ybn,N . Suppose that the following conditions hold.
(a) There exists a non-random element y of H such that

plim ||YN − y|| = 0. (3.14)


N →∞

(b) There exist a sequence {xm }∞


m=1 of non-random elements of H and a

sequence {ρm }m=1 of positive numbers such that

X
n
plim ρm ||Xm,N − xm || = 0 (3.15)
N→∞ m=1

and
° n °
°X °
° °
liminf ° ρm xm ° > 0. (3.16)
n→∞ ° °
m=1

Then plimN→∞ ||Ybn,N − yb|| = 0 and plimN→∞ ||Un,N − u|| = 0, where yb is the
projection of y on span({xm }∞ b is the residual involved.
m=1 ) and u = y − y
3.6. APPENDIX: PROOFS 33

3.6 Appendix: Proofs


3.6.1 Theorem 3.1
Recall that ”subspace” means a sub-Hilbert space. Thus, S is a Hilbert
space.
Pick a sequence zn ∈ S such that

||y − zn || ≤ ||y − yb|| + n−1 . (3.17)

This is always possible because otherwise ||y − z|| > ||y − yb|| + n−1 for all
z ∈ S so that inf z∈S ||y − z|| ≥ ||y − yb|| + n−1 . Then

lim ||y − zn ||2 = ||y − yb||2 = δ. (3.18)


n→∞

say. The first step is to show that zn is a Cauchy sequence. Observe that

||zn − zm ||2 = ||(zn − y) − (zm − y)||2


= ||zn − y||2 − 2 hzn − y, zm − yi + ||zm − y||2

and

4.||0.5(zn + zm ) − y||2 = ||(zn − y) + (zm − y)||2


= ||zn − y||2 + 2 hzn − y, zm − yi + ||zm − y||2

Adding these two equation up yields

||zn − zm ||2 = 2||zn − y||2 + 2||zm − y||2 − 4.||0.5(zn + zm ) − y||2 (3.19)

Because 0.5(zn + zm ) ∈ S, it follows that ||0.5(zn + zm ) − y||2 ≥ δ 2 , whereas


2 2
by (3.17) and (3.18), ||zn − y||2 ≤ (δ + n−1 ) and ||zm − y||2 ≤ (δ + m−1 ) .
Therefore, it follows from (3.19) that
¡ ¢2 ¡ ¢2
||zn − zm ||2 ≤ 2 δ + n−1 + 2 δ + m−1 − 4δ 2
= 4δ/n + 2n−2 + 4δ/m + 2m−2 .

Thus, zn is a Cauchy sequence in S and therefore takes a limit yb in S.


The next step is to show that for all z ∈ S, hy − yb, zi = 0, as follows.
Note that for any real scalar c, yb + c.z ∈ S and therefore

||y − yb||2 ≤ ||y − yb − c.z||2 = ||y − yb||2 − 2c. hy − yb, zi + c2 ||z||2


34 CHAPTER 3. PROJECTIONS

The right-hand side is minimal for c = hy − yb, zi /||z||2 , hence

(hy − yb, zi)2


0≤−
||z||2

and thus hy − yb, zi = 0.


Note that this argument only applies if the Hilbert space H is real. If H
is complex this orthogonality proof can be adapted similar to the proof of
Theorem 2.1.
Finally, we need to show that yb is unique. Suppose that there exists an-
other projection ye ∈ S. Then also hy − ye, zi = 0, and thus hy − ye, zi −
hy − yb, zi = hb y − ye, zi = 0. But z = y − ye ∈ S so that ||b y − ye||2 =
y − ye, yb − yei = 0. Consequently, yb is unique.
hb

3.6.2 Lemma 3.1


Let Mn = span({xk }nk=1 ) and M∞ = span({xk }∞ ∞
k=1 ) = ∪n=1 Mn . If z ∈
∪∞n=1 Mn then there exists an n0 such that z ∈ Mn0 , hence for n ≥ n0 ,
zbn = z and thus limn→∞ ||z − zbn || = 0. Now let z ∈ M∞ \(∪∞ n=1 Mn ). Since

M∞ = ∪n=1 Mn is closed and Mn ⊂ Mn+1 , for each n there exists an
zn ∈ Mn such that limn→∞ ||z − zn ||2 = 0, hence for n → ∞, ||z − zbn ||2 ≤
||z − zn ||2 → 0.

3.6.3 Theorem 3.2


Adopting the notation in the proof of Lemma 3.1, we may without loss
of generality assume that zb ∈ M∞ \(∪∞ n=1 Mn ), as otherwise the result of
Theorem 3.2 holds trivially. Since M∞ is closed this assumption implies
that for each n we can select a zn ∈ Mn such that

lim ||b
z − zn || = 0. (3.20)
n→∞

Let ||z − zb|| = δ and ||z − zbn || = δn , and note that δn ≥ δ. Since

δn2 = ||z − zbn ||2 ≤ ||z − zn ||2 = ||z − zb + zb − zn ||2


= ||z − zb||2 + ||b
z − zn ||2 + 2 hz − zb, zb − zn i
= δ 2 + ||b
z − zn ||2
3.6. APPENDIX: PROOFS 35

it follows from (3.20) that


lim δn = δ. (3.21)
n→∞
Recall that z = zb + u, where hu, xi = 0 for all x ∈ M∞ . Hence
z − zbn ||2 = ||z − zbn − u||2 = ||z − zbn ||2 + ||u||2 − 2 hz − zbn , ui
||b
= ||z − zbn ||2 + ||u||2 − 2 hz, ui = δn2 − δ2 (3.22)
where the last equality follows from hz, ui − hu, ui = hb
z , ui = 0 and hu, ui =
||u||2 = δ 2 . The theorem now follows from (3.21) and (3.22).

3.6.4 Theorem 3.3


Due to the orthonormality condition (3.7), the projection zbn of z on Mn =
span({xj }nj=1 ) takes the form
X
n
zbn = θj xj , where θj = hz, xj i . (3.23)
j=1

Moreover, denoting un = z − zbn , it follows from (3.7) and (3.23) that


° °2
° Xn ° X
n X
n X
n
2 ° ° 2
||un || = °z − θj xj ° = ||z|| − 2 θj hz, xj i + θj θi hxj , xi i
° °
j=1 j=1 j=1 i=1
X
n
= ||z||2 − θj2 ≥ 0 (3.24)
j=1
P P
hence nj=1 θj2 ≤ ||z||2 for all n and thus ∞ 2
j=1 θj < ∞. Finally, it follows
from Theorem 3.2 that
° °2
° Xn °
° °
lim °zb − z − zbn ||2 = 0
θj xj ° = lim ||b
n→∞ ° ° n→∞
j=1
P∞
so that we can write zb = j=1 θj xj .

3.6.5 Lemma 3.2


Let x be an arbitrary element of a subspace S of an Hilbert space H and
let yn be a Cauchy sequence in S ⊥ . Then there exists an y ∈ H such that
limn→∞ ||y − yn || = 0. Since hx, yn i = 0 we have hx, yi = hx, y − yn i . It
follows now from the Cauchy-Schwarz inequality that | hx, yi | = | hx, y − yn i |
≤ ||x||.||y − yn || → 0. Hence y ∈ S ⊥ .
36 CHAPTER 3. PROJECTIONS

3.6.6 Theorem 3.4


Denote Sn = span({xk }∞ k=n ). Project each xk on Sk+1 , so that xk = x bk +
uk with projection x bk ∈ Sk+1 and residual uk . Recall that by the regularity
condition, ||uk || > 0, hence ek = uk /||uk || is well defined. It is not hard to
verify that the residuals uk are orthogonal, so that the ek ’s are orthonormal.
Next, denote

Un = span (e1 , ..., en ) = span (u1 , ..., un ) ,

and let Un⊥ be the orthogonal complement of Un . Note that



Un+1 ⊂ Un⊥ . (3.25)

To see this, let z ∈ Un+1 . Then for all x ∈ Un+1 , hz, xi = 0, and because
obviously Un ⊂ Un+1 , it follows that also hz, xi = 0 for all x ∈ Un . Hence,
z ∈ Un⊥ .
As before, let Mn = span({xk }nk=1 ).
The theorem under review will be proved in six steps:

Step 1. First I will show that


¡ ¢
Mn ⊂ span Un , Un⊥ ∩ S2 . (3.26)

Proof. Let z ∈ Mn be arbitrary. Recall that z takes the form z =


P n
bk + uk = x
k=1 ck xk . Substituting xk = x bk + ||uk ||ek we can write z as
X
n X
n X
n
z = ck (b
xk + uk ) = ck uk + bk
ck x
k=1 k=1 k=1
X
n X
n
= ck ||uk ||ek + bk
ck x
k=1 k=1

Note that
X
n
bk ∈ S2
ck x (3.27)
k=1

because xbk ∈ Sk+1 P⊂ S2 .


n
Pn Next, project bk on Un . This projection takes the form pbn =
k=1 ck x
k=1 dk ek with residual wn+1 ∈ S2 . The latter follows from (3.27). But since
3.6. APPENDIX: PROOFS 37

wn+1 is a residual of a projection on Un we also have hek , wn+1 i = 0 for


k = 1, ..., n, hence wn+1 ∈ Un⊥ . Thus,
wn+1 ∈ Un⊥ ∩ S2 .
Denoting αk = ck ||uk || + dk , we can now write
X
n
z= αk ek + wn+1 , where wn+1 ∈ Un⊥ ∩ S2 .
k=1

Therefore, (3.26) holds.

Step 2. I will now show that


¡ ¢ ¡ ¢
span Un , Un⊥ ∩ S2 = span Un , Un⊥ ∩ Sn+1 . (3.28)

Proof. Denote
Sk,m = span({xj }m
j=k )

for m ≥ k and let z ∈ Un⊥ ∩ S2,m for some m ≥ 2. Consider first the case
m > n. Since z ∈ S2,m there exists constants ck such that
X
m X
n X
m
z = ck xk = ck (b
xk + uk ) + ck xk
k=2 k=2 k=n+1
Xn X
n X
m
= ck kuk k ek + bk +
ck x ck xk .
k=2 k=2 k=n+1

Moreover, since z ∈ Un⊥ it follows that hz, ek i = 0 for k = 1, ..., n. In


particular,
X
n X
m
0 = hz, e2 i = c2 ku2 k + ck hb
xk , e2 i + ck hxk , e2 i
k=2 k=n+1
= c2 ku2 k
P P
because nk=2 ck x
bk ∈ S3 , m k=n+1 ck xk ∈ Sn+1 , and e2 is orthogonal to S3
and Sn+1 . Hence c2 = 0 and thus
X
n X
n X
m
z= ck kuk k ek + bk +
ck x ck xk .
k=3 k=3 k=n+1
38 CHAPTER 3. PROJECTIONS

It follows now similarly that ck = 0 for k = 3, ...., n, hence


X
m
z= ck xk ∈ Sn+1,m .
k=n+1

Because z ∈ Un⊥ as well, it follows now that

z ∈ Un⊥ ∩ Sn+1,m ,

which implies
Un⊥ ∩ S2,m ⊂ Un⊥ ∩ Sn+1,m
because z ∈ Un⊥ ∩ S2,m was arbitrary. However, Sn+1,m ⊂ S2,m and therefore

Un⊥ ∩ Sn+1,m ⊂ Un⊥ ∩ S2,m ,

so that
Un⊥ ∩ S2,m = Un⊥ ∩ Sn+1,m for m > n.
This result implies that
¡ ¢ ¡ ∞ ¢
Un⊥ ∩ ∪∞ ⊥
m=n+1 S2,m = Un ∩ ∪m=n+1 Sn+1,m (3.29)

In the case m ≤ n, z ∈ Un⊥ ∩ S2,m implies that z = 0, as can be straight-


forwardly verified from the above argument, so that Un⊥ ∩ S2,m = {0} for
m = 2, 3, ..., n. Since Hilbert spaces are vector spaces and therefore always
contain the null element it follows that
¡ ∞ ¢
∪∞ ⊥ ⊥
m=2 Un ∩ S2,m = {0} ∪ ∪m=n+1 Un ∩ S2,m
= ∪∞ ⊥
m=n+1 Un ∩ S2,m ,

hence ¡ ∞ ¢
Un⊥ ∩ (∪∞ ⊥
m=2 S2,m ) = Un ∩ ∪m=n+1 S2,m . (3.30)
Since by Definition 2.10,

S2 = ∪∞ ∞
m=2 S2,m , Sn+1 = ∪m=n+1 Sn+1,m

it follows now from (3.30) that

Un⊥ ∩ S2 = Un⊥ ∩ ∪∞
m=2 S2,m
⊥ ⊥
= Un ∩ ∪∞m=n+1 Sn+1,m = Un ∩ Sn+1
3.6. APPENDIX: PROOFS 39

which implies that (3.28) holds.


¡ ¢
Step 3. Denote Rn = span Un , Un⊥ ∩ Sn+1 . Then

S1 = ∪∞
n=1 Rn . (3.31)

Proof. Combining (3.26) and (3.28) yields Mn ⊂ Rn , hence

S1 = ∪∞ ∞
n=1 Mn ⊂ ∪n=1 Rn , (3.32)

where the equality follows from Definition 2.10. However, we also have Rn ⊂
S1 , as is not hard to verify, hence

∪∞
n=1 Rn ⊂ S1 . (3.33)

Thus, the result (3.31) follows from (3.32) and (3.33).

bn be the projection of x on Rn . Then


Step 4. For an x ∈ S1 , let x
X
n
bn =
x αj ej + wn+1 (3.34)
j=1

where αj = hx, ej i and wn+1 is the projection of x on Un⊥ ∩ Sn+1 . Moreover,

X

αj2 < ∞. (3.35)
j=1

Furthermore, ° °
° X
n °
° °
lim °x − αj ej − wn+1 ° = 0. (3.36)
n→∞ ° °
j=1
P
Proof. By the definition of Rn and by Definition 3.3, x bn = nj=1 θj ej + w
for some constants θj and a w ∈ Un⊥ ∩ Sn+1 . To determine the θj ’s and w,
note that
° °2
° Xn ° X
n X
n
° ° 2
°x − θ e
j j − w ° = ||x − w|| − 2 θj hej , xi + 2 θj hej , wi
° °
j=1 j=1 j=1
40 CHAPTER 3. PROJECTIONS
° n °2
°X °
° °
+° θj ej °
° °
j=1
X
n X
n
2
= ||x − w|| − 2 θj hej , xi + θj2
j=1 j=1

because w ∈ Un⊥ ∩ Sn+1 ⊂ Un⊥ implies hej , wi = 0 and


° n °2
°X ° X
n X n X
n X
n
° ° 2
° θj ej ° = θj θi hej , ei i = θj hej , ej i = θj2 .
° °
j=1 j=1 i=1 j=1 j=1

Thus
° °2
° Xn °
° °
bn ||2
||x − x = inf °x − θj ej − w°
θ1 ,...,θn ,w∈Un ∩Sn+1 °
⊥ °
j=1
à !
X n X
n
= inf ||x − w||2 − 2 θj hej , xi + θj2
θ1 ,...,θn ,w∈Un⊥ ∩Sn+1
j=1 j=1
X
n
= inf ||x − w||2 − αj2
w∈Un⊥ ∩Sn+1
j=1
X
n
= ||x − wn+1 ||2 − αj2 (3.37)
j=1

where αj = hx, ej i and wn+1 is the projection of x on Un⊥ ∩ Sn+1 .


This result implies that for all n,
X
n
αj2 ≤ ||x − wn+1 ||2 ≤ ||x||2 (3.38)
j=1

so that (3.35) holds.


b be the projection of x on ∪∞
Finally, to prove (3.36), let x n=1 Rn . Then
it follows from Theorem 3.2 that limn→∞ ||b xn − xb|| = 0. But (3.31) implies
b ∈ S1 , hence x = x
x b, so that limn→∞ ||b
xn − x|| = 0.
Pn
Step 5. Let zn = j=1 αj ej . Then
lim kz − zn k = 0, where z ∈ U∞ . (3.39)
n→∞
3.6. APPENDIX: PROOFS 41

Proof. This follows from the fact that zn is a Cauchy sequence in U∞ =


span({ek }∞
k=1 ) because
° °2
° max(m,n)
X °
° °
kzn − zm k = °
2
° αj ej
°
°
°j=min(m,n)+1 °
max(m,n)
X X

= αj2 ≤ αj2
j=min(m,n)+1 j=min(m,n)+1
→ 0
P∞
as min(m, n) → ∞, where the latter is due to j=1 αj2 < ∞.


Step 6. There exists a w ∈ U∞ ∩ S∞ such that
lim ||wn+1 − w|| = 0. (3.40)
n→∞

Proof. Recall from Step 4 that


wn+1 ∈ Un⊥ ∩ Sn+1 .
Moreover, it follows from (3.25) and the definition of Sn+1 that for an arbi-
trary k ≥ 1,
Un⊥ ∩ Sn+1 ⊂ Uk⊥ ∩ Sk+1 for n ≥ k
hence
wn+1 ∈ Uk⊥ ∩ Sk+1 for n ≥ k.
Furthermore for n ≥ k, wn+1 is a Cauchy sequence in Uk⊥ ∩ Sk+1 because
||wn+1 − wm+1 || = ||b
xn − zn − xbm + zm ||
≤ bm || + ||zn − zm ||
xn − x
||b
≤ xn − x|| + ||b
||b xm − x|| + ||zn − zm ||
→ 0
as min(m, n) → ∞. Thus, there exists a w ∈ Uk⊥ ∩Sk+1 such that (3.40) holds.
Since k was arbitrary we have w ∈ ∩∞ ⊥ ⊥ ∞
k=1 Uk = U∞ and w ∈ ∩k=1 Sk+1 = S∞ ,
hence

w ∈ U∞ ∩ S∞ .
42 CHAPTER 3. PROJECTIONS

This completes the proof of Step 6.

The theorem now follows from (3.35), (3.39), (3.40) and the fact that
⊥ ⊥
w ∈ U∞ ∩ S∞ ⊂ U∞ , which implies that hw, ek i = 0 for k ∈ N.

3.6.7 Theorem 3.5


r h i
et / E U
Recall that Ut = U et = Xt − X
et2 , where U bt with X
bt the projection of

Xt on span({Xt−j }∞ e
j=1 ). The uncorrelatedness of the Ut ’s follows from The-
orem 3.4, but we still need to show that E[Uet ] = 0 and E[U et2 ] = σ 2 for all
t.

Proof of E[U et ] = 0
Let Xbt,n be the projection of Xt on span({Xt−j }nj=1 ). Then X
bt,n takes the
form
X
n
Xbt,n = βj,n Xt−j ,
j=1

where the βj,n ’s do not depend on t. The latter follows from the fact that
the βj,n ’s are the solutions of the normal equations
X
n
βj,n γ(i − j) = γ(i), i = 1, 2, ..., n,
j=1
h i
where γ(i) = E [Xt Xt−i ] is the covariance function of Xt . Hence E Xbt,n =
0.
It follows from Theorem 3.2 that
° °2 ∙³ ´2 ¸
°b b ° b b
lim °Xt,n − Xt ° = lim E Xt,n − Xt =0 (3.41)
n→∞ n→∞

h i
so that by Liapounov’s inequality and E Xbt,n = 0,
¯ ¯ ¯ ¯ h¯ ¯i
¯ b ¯ ¯ b b ¯ ¯b b ¯
lim ¯E[Xt ¯ =
] lim ¯E[X t − X t,n ¯
] ≤ lim E ¯ t
X − X t,n ¯
n→∞ n→∞ n→∞
s ∙³ ´2 ¸
≤ b
lim E Xt,n − Xt b = 0.
n→∞
3.6. APPENDIX: PROOFS 43

bt ] = 0 and therefore E[U


Thus E[X et ] = E[Xt − X
bt ] = 0.

Proof of E[U et2 ] = σ 2


et,n = Xt − X
Let U bt,n . It follows from (3.41) that
∙³ ´2 ¸ ∙³ ´2 ¸
e
lim E Ut − Ut,n e b b
= lim E Xt,n − Xt =0 (3.42)
n→∞ n→∞

Moreover,
⎡Ã !2 ⎤
h i ° °2 X
n
et,n ° bt,n °
E U 2
= °Xt − X ° = E ⎣ Xt − βj,n Xt−j ⎦
j=1

X
n X
n X
n
= γ (0) − 2 βj,n γ(j) + βj,n βi,n γ(i − j)
j=1 j=1 i=1

= σn2
say, which does not depend on t. Furthermore, note that σn2 is non-increasing
in n, so that
lim σn2 = σ 2
n→∞
exists, and that
∙³ ´2 ¸ ° °2 ° °2
e e °b b ° °b e °
E Ut − Ut,n = °Xt,n − Xt ° = °Xt,n − Xt + Ut °
° °2 D E
°b ° b e e 2
= °X t,n − X t ° + 2 X t,n − X t , Ut + ||Ut ||
° ° D E
° e °2 e e 2
= °U t,n ° − 2 X t , U t + ||Ut ||
° ° D E
° e °2 b e e e 2
= °U t,n ° − 2 Xt + Ut , Ut + ||Ut ||
° ° D E
° e °2 et , U et + ||U et ||2
= °Ut,n ° − 2 U
° °
° e °2 et ||2
= °Ut,n ° − ||U
h i h i
= E U et,n
2
−E U et2 .

Thus, ∙³
h i ´2 ¸
E et2
U = σn2 −E U et − U
et,n → σ2.
44 CHAPTER 3. PROJECTIONS

Proof of (3.10), (3.11) and (3.13)


The result of Theorem 3.4 can now be translated as
° °
° Xn °
° °
lim °Xt − αj Ut−j − Wt ° = 0, (3.43)
n→∞ ° °
j=0

where Ut is a zero-mean uncorrelated covariance stationary


P∞ 2 process with unit
variance, and αk = hXt , Ut−k i = E [Xt Ut−k ] with k=0 αk < ∞.
We still need to prove that the αk ’s do not depend
Pn on t, as follows. Recall
e 2 2 e
from the proof of E[Ut ] = σ that Ut,n = Xt − j=1 βj,n Xt−j , so that
h i X
n
e
E Xt+k Ut,n = γ (k) − βj,n γ(k + j),
j=1

which does not depend on t. Moreover, by the Cauchy-Schwarz inequality


and (3.42),
¯ h ³ ´i¯2 ∙³ ´2 ¸
¯ e e ¯ e e
lim ¯E Xt+k Ut,n − Ut ¯ ≤ γ (0) lim E Ut,n − Ut = 0.
n→∞ n→∞

h i h i
Thus E Xt+k U et = limn→∞ E Xt+k U et,n . Since the latter does not depend
h i
on t, neither does αk = E [Xt+k Ut ] = E Xt+k U et || .
et /||U
The results (3.11) and (3.13) follow straightforwardly from Theorem 3.4.

Proof of (3.9)
The result (3.43) implies, by Chebyshev’s inequality, that
X
n
Xt = plim αj Ut−j + Wt . (3.44)
n→∞
j=0

Recall that convergence in probability for n → ∞ is equivalent to a.s.


convergence along a further subsequence km of an arbitrary subsequence of
n. See for example Bierens (2004, Theorem 6.B.3, p. 168). Thus for such a
subsequence km ,
X
km
a.s.
αj Ut−j → Xt − Wt (3.45)
j=0
3.6. APPENDIX: PROOFS 45

as m → ∞, and the same holds for any further subsequence of km .


Without loss of generality we may choose k0 = 0. Then for each n > 0 we
can find an mn such that

kmn−1 < n ≤ kmn . (3.46)

Moreover, (3.45) implies that


kmn
X a.s.
αj Ut−j → Xt − Wt as n → ∞. (3.47)
j=0

Due to (3.46),
⎡Ã !2 ⎤ ⎡Ã !2 ⎤
kmn kmn
X∞ X X
n X
∞ X
E⎣ αj Ut−j − αj Ut−j ⎦ = E⎣ αj Ut−j ⎦
n=1 j=0 j=0 n=1 j=n+1
kmn
X
∞ X X

≤ αj2 ≤ αj2 < ∞,
n=1 j=kmn−1 +1 j=0

so that by Chebyshev’s inequality, for arbitrary ε > 0,


"¯ km ¯ #
X∞ ¯X n Xn ¯
¯ ¯
Pr ¯ αj Ut−j − αj Ut−j ¯ > ε < ∞.
n=0
¯ ¯
j=0 j=0

This result implies, by the Borel-Cantelli lemma,7 that


kmn
X X
n
a.s.
αj Ut−j − αj Ut−j → 0 as n → ∞. (3.48)
j=0 j=0

Combining (3.47) and (3.48) it follow now that


X
n
a.s.
αj Ut−j → Xt − Wt as n → ∞. (3.49)
j=0
P Pn
Since ∞ j=0 αj Ut−j is defined as limn→∞ j=0 αj Ut−j , (3.9) is equivalent to
(3.49).
7
See for example Bierens (2004, Theorem 6.B.2, p. 168).
46 CHAPTER 3. PROJECTIONS

The zero-mean covariance stationarity of Wt


It follows now trivially from (3.9) that E [Wt ] = 0. Moreover, it is left as an
exercise to show that for m ≥ 0,
X

E [Wt Wt−m ] = γ (m) − αj+m αj . (3.50)
j=0

Proof of (3.12)
Finally, Wt ∈ ∩n span({Xn−j }∞ ∞
j=0 ) implies that Wt ∈ span({Xt−j }j=1 ), hence
the projection of Wt on span({Xt−j }∞ j=1 ) is Wt itself. Since by (3.9),
¡ ¢
span({Xt−j }∞ ∞ ∞
j=1 ) = span span({Ut−j }j=1 ), span({Wt−j }j=1 )

and the projection of Wt on span({Ut−j }∞ j=1 ) is zero, it follows that the pro-
jection of Wt on span({Wt−j }∞
j=1 ) is Wt itself, which proves (3.12).

3.6.8 Theorem 3.7


Note that ||Ybn,N − yb|| = ||(u − Un,N ) − (YN − y)|| ≤ ||Un,N − u|| + ||YN − y||,
hence
||Ybn,N − yb|| ≤ ||Un,N − u|| + op (1),
where the op (1) term follows from condition (a). Therefore it suffices to prove
||Un,N − u|| = op (1), as follows.
en,N =
Let Yen,N be the projection of y on span({Xm,N }nm=1 ) , with residual U
y − Yen,N , and let ybn be the projection of y on span({xm }m=1 ) , with residual
n

un = y − ybn . Then by the triangular inequality,


en,N || + ||un − U
||Un,N − un || ≤ ||Un,N − U en,N ||.

It will be shown that


en,N || = op (1)
||Un,N − U (3.51)

and
yn − Yen,N || = ||un − U
||b en,N || = op (1). (3.52)
yn − yb|| = 0 and thus limn→∞ ||un − u|| = 0, the result of the
Since limn→∞ ||b
theorem under review then follows from (3.51) and (3.52).
3.6. APPENDIX: PROOFS 47

Proof of (3.51)
Denote the angle between two elements x and y of H by ϕ(x, y). Recall that
³ ´
sin2 ϕ(YN , Ybn,N ) = ||Un,N ||2 /||YN ||2
³ ´
sin2 ϕ(y, Yen,N ) = ||Uen,N ||2 /||y||2
³ ´
cos ϕ(YN , Yn,N ) = ||Ybn,N ||/||YN ||
b
³ ´
cos ϕ(y, Yen,N ) = ||Yen,N ||/||y||.

Using these formulas we can write


D E
||Un,N e 2 2 e 2 e
− Un,N || = ||Un,N || + ||Un,N || − 2 Un,N , Un,N
³ ´ ³ ´
= ||YN ||2 sin2 ϕ(YN , Ybn,N ) + ||y||2 sin2 ϕ(y, Yen,N )
D E
−2 Un,N , U en,N

= ||YN ||2 + ||y||2


³ ´ ³ ´
−||YN ||2 cos2 ϕ(YN , Ybn,N ) − ||y||2 cos2 ϕ(y, Yen,N )
D E
−2 Un,N , U en,N
³ ´
2 2 2 b
= ||YN − y|| − ||YN || cos ϕ(YN , Yn,N )
³ ´ D E
−||y||2 cos2 ϕ(y, Yen,N ) + 2 hYN , yi − 2 Un,N , Uen,N

and
D E
en,N
Un,N , U en,N + Yen,N i = hUn,N , yi
= hUn,N , U
= hUn,N + Ybn,N , yi − hYbn,N , yi
= hYN , yi − hYbn,N , yi
³ ´
= hYN , yi − cos ϕ(y, Ybn,N ) ||Ybn,N ||.||y||
³ ´ ³ ´
= hYN , yi − cos ϕ(y, Ybn,N ) cos ϕ(YN , Ybn,N ) ||y||.||YN ||
³ ´ ³ ´
e b
≥ hYN , yi − cos ϕ(y, Yn,N ) cos ϕ(YN , Yn,N ) ||y||.||YN ||
48 CHAPTER 3. PROJECTIONS
³ ´ ³ ´
where the inequality follows from cos ϕ(y, Ybn,N ) ≤ cos ϕ(y, Yen,N ) . Thus

en,N ||2 ≤ ||YN − y||2


||Un,N − U
³ ³ ´ ³ ´´2
− ||YN || cos ϕ(YN , Ybn,N ) − ||y|| cos ϕ(y, Yen,N )
≤ ||YN − y||2 = op (1)
where the op (1) term is due to condition (3.14). This proves (3.51).

Proof of (3.52)
Let r1 = x1 and for m ≥ 2, let rm be the residual of the projection of xm
on span(x1 , ..., xm−1 ). Denote em = ||rm ||−1 rm if ||rm || > 0 and em = 0 if
||rm || = 0. Similarly, let R1,N = X1,N and for m = 2, ..., n, let Rm,N be
the residual of the projection of Xm,N on span(X1,N , ..., Xm−1,N ). Denote
ebm,N = ||Rm,N ||−1 Rm,N if ||Rm,N || > 0, and ebm,N = 0 if ||Rm,N || = 0. Then
we can write
Xn X∞
2
ybn = αm em , where αm = hy, em i and αm < ∞ (3.53)
m=1 m=1
Xn
Yen,N = b m,N ebm,N , where α
α b m,N = hy, b
em,N i . (3.54)
m=1

D It followsE fromD the trivial


E equalities
D yn E− Yen,N ||2 = ||DYen,N ||2 +E||b
||b yn ||2 −
2 ybn , Yen,N and ybn , Yen,N = ybn , y − Uen,N = ||b yn ||2 − ybn , Uen,N that
D E
||b e 2 e 2 2 e
yn || + 2 ybn , Un,N .
yn − Yn,N || = ||Yn,N || − ||b

Moreover, using the Cauchy-Schwarz inequality and the fact that ||U en,N || ≤
||y||, it follows that
¯* n +¯ ¯* n +¯
¯D E¯ ¯ X ¯ ¯ X ¯
¯ e ¯ ¯ e ¯ ¯ e ¯
¯ ybn , Un,N ¯ = ¯ αm em , Un,N ¯ = ¯ αm (em − ebm,N ), Un,N ¯
¯ m=1 ¯ ¯ m=1 ¯
° n °
°X °
e ° °
≤ ||Un,N ||. ° αm (em − ebm,N )°
° °
° n m=1 °
°X °
° °
≤ ||y||. ° αm (em − ebm,N )°
°m=1 °
3.6. APPENDIX: PROOFS 49
° k ° v
°X ° u X
° ° u ∞
≤ ||y||. ° α (e − ebm,N )° + 2||y||t 2
αm
°m=1 m m °
m=k+1

qP
∞ 2
Given an arbitrary ε > 0 we can choose k so large that 2||y|| m=k+1 αm <
°P °
° °
ε, and for this k, ° km=1 αm (em − b em,N )° = op (1), as is easy to verify from
condition (3.15). Consequently,
D E
en,N = op (1)
ybn , U

and thus
yn − Yen,N ||2 = ||Yen,N ||2 − ||b
||b yn ||2 + op (1). (3.55)
The next step is to show that

||Yen,N || ≤ ||b
yn || + op (1), (3.56)

as follows. Note that


° °2
° X n °
en,N || = ° °
||U inf °y − βm Xm,N °
β1 ,...,βn ° °
m=1
° °2
° Xn °
° °
= inf inf °y − λ ξm Xm,N °
(ξ1 ,...,ξn ) ∈Xm=1 [−ρm ,ρm ] λ °
0 n °
m=1
( µ P ¶2 )
2 hy, nm=1 ξm Xm,N i
= inf ||y|| − P
(ξ1 ,...,ξn )0 ∈Xnm=1 [−ρm ,ρm ] k nm=1 ξm Xm,N k
P 2
2 (hy, nm=1 ξm Xm,N i)
= ||y|| − sup P 2
(ξ1 ,...,ξn )0 ∈Xn
m=1 [−ρm ,ρm ] k nm=1 ξm Xm,N k

hence P
hy, nm=1 ξm Xm,N i
||Yen,N || = sup P (3.57)
(ξ1 ,...,ξn )0 ∈Xn
m=1 [−ρm ,ρm ]
k nm=1 ξm Xm,N k
and similarly,
P
hy, nm=1 ξm xm i
yn || =
||b sup P (3.58)
(ξ1 ,...,ξn )0 ∈Xn
m=1 [−ρm ,ρm ]
k nm=1 ξm xm k
50 CHAPTER 3. PROJECTIONS

Note that by condition (3.16) at least one xm is non-zero, so that (3.57) and
(3.58) are well-defined for sufficiently large n.
Since the ratios in (3.57) and (3.58) are scale-invariant, we may without
loss of generality impose the normalization
° n ° ° n °
°X ° 1° °
° ° °X °
° ξm xm ° = Mn = ° ρm xm ° , (3.59)
°m=1 ° 2 °m=1 °

for example. Note that (3.59) is compatible with (ξ1 , ..., ξn )0 ∈ Xnm=1 [−ρm , ρm ].
Thus, denoting
( ° n ° )
°X °
° °
Ξn = (ξ1 , ..., ξn )0 ∈ Xnm=1 [−ρm , ρm ] : ° ξ x ° = Mn ,
°m=1 m m °

the expressions (3.57) and (3.58) are equivalent to


Pn
hy, ξ X i
||Yen,N || = sup Pn m=1 m m,N (3.60)
(ξ1 ,...,ξn )0 ∈Ξn k m=1 ξm Xm,N k

and P
hy, nm=1 ξm xm i
||b
yn || = sup P , (3.61)
(ξ1 ,...,ξn )0 ∈Ξn k nm=1 ξm xm k
respectively.
Finally, observe from (3.59), (3.60) and (3.61) that
½ Pn P
hy, m=1 ξm Xm,N i k nm=1 ξm Xm,N k
yn || =
||b sup P × Pn
(ξ1 ,...,ξn )0 ∈Ξn k nm=1 ξm Xm,N k k m=1 ξm xm k
Pn ¾
hy, m=1 ξm (Xm,N − xm )i
− P
k nm=1 ξm xm k
½ Pn P
hy, m=1 ξm Xm,N i k nm=1 ξm Xm,N k
= sup P ×
(ξ1 ,...,ξn )0 ∈Ξn k nm=1 ξm Xm,N k Mn
Pn ¾
hy, m=1 ξm (Xm,N − xm )i

Mn
µ Pn ¶ P
m=1 ρm ||Xm,N − xm || hy, nm=1 ξm Xm,N i
≥ 1− sup Pn
Mn (ξ1 ,...,ξn )0 ∈Ξn k m=1 ξm Xm,N k
Pn
ρm ||Xm,N − xm ||
−||y|| m=1
Mn
3.6. APPENDIX: PROOFS 51
µ Pn ¶ Pn
ρ ||X − x || ρm ||Xm,N − xm ||
||Yen,N || − ||y|| m=1
m m,N m
= 1 − m=1
M Mn
Pnn
ρm ||Xm,N − xm ||
≥ ||Yen,N || − 2||y|| m=1 , (3.62)
Mn

where the last inequality follows from ||Yen,N || ≤ ||y||. Since by condition
(3.16), liminfn→∞ Mn > 0, it follows from (3.15) and (3.62) that (3.56) holds.
The latter together with (3.55) imply (3.52).
52 CHAPTER 3. PROJECTIONS
Chapter 4

Orthogonal polynomials

4.1 Introduction
Let w(x) be a non-negative Borel measurable real-valued function on R sat-
isfying
Z ∞
|x|k w(x)dx ∈ (0, ∞) for k ∈ N0
−∞

where the integral involved is the Lebesgue integral. Without loss of general-
ity we may assume that w is a density function with finite absolute moments
of any order. Let

X
k
pk (x|w) = αk,j xj , αk,k = 1, k ∈ N0 (4.1)
j=0

be a sequence of polynomials in x ∈ R such that


Z ∞
pk (x|w)pm (x|w)w(x)dx = 0 if k 6= m. (4.2)
−∞

In words, the polynomials pk (x|w) are orthogonal with respect to the weight
function w(x).
Defining
pk (x|w)
pk (x|w) = qR (4.3)
∞ 2
p (y|w) w(y)dy
−∞ k

53
54 CHAPTER 4. ORTHOGONAL POLYNOMIALS

yields a sequence of orthonormal polynomials w.r.t. w(x):


Z ∞ ½
0 if k 6= m,
pk (x|w)pm (x|w)w(x)dx = (4.4)
−∞ 1 if k = m.

This sequence is uniquely determined by w(x), except for signs. In other


words, |pk (x|w)| is unique. To show this, suppose that there exists another
sequence p∗k (x|w) of orthonormal polynomials w.r.t.
P w(x). Since p∗k (x|w) is a
polynomial of order k, we can write p∗k (x|w) = km=0 βm,k pm (x|w). Similarly,
P
we can write pk (x|w) = km=0 αm,.k p∗m (x|w). Then for j < k,
Z ∞ j
X Z ∞
p∗k (x|w)pj (x|w)w(x)dx = αm,j p∗k (x|w)p∗m (x|w)w(x)dx
−∞ m=0 −∞
= 0

and
Z ∞ X
k Z ∞
p∗k (x|w)pj (x|w)w(x)dx = βm,k pm (x|w)pj (x|w)w(x)dx
−∞ m=0 −∞
Z ∞
= βj,k pj (x|w)2 w(x)dx = βj,k ,
−∞

hence βj,.k = 0 for j < k and thus

p∗k (x|w) = βk,k pk (x|w).

Moreover, by normality,
Z ∞ Z ∞
∗ 2 2
1= pk (x|w) w(x)dx = βk,k pk (x|w)2 w(x)dx = βk,k
2
,
−∞ −∞

so that p∗k (x|w) = ±pk (x|w). Consequently, |pk (x|w)| is unique.


The reason for considering orthonormal polynomials is the following.

Theorem 4.1. Let w(x) be a density function with support (a, b), −∞ ≤
a < b ≤ ∞, satisfying the moment conditions
Z ∞
|x|k w(x)dx < ∞ (4.5)
−∞
4.1. INTRODUCTION 55

for k ∈ N0 . Denote by L2 (w) be the Hilbert space of Borel measurable real


Rb
functions f on (a, b) satisfying a f (x)2 w(x)dx < ∞, with inner product
Rb p
hf, gi = a f (x)g(x)w(x)dx and associated norm kf k = hf, f i and metric
||f − g||. For an arbitrary function f ∈ L2 (w) , let
X
n
fn (x) = γk pk (x|w),
k=0

where Z b
γk = hf, pk i = f (x)pk (x|w)w(x)dx. (4.6)
a
Then
lim kf − fn k = 0. (4.7)
n→∞

This result implies that every function f ∈ L2 (w) can be written as


X

f (x) = γk pk (x|w) a.e. on (a, b). (4.8)
k=0

Note that condition (4.5) holds trivially if the support (a, b) of w(x) is
bounded. However, as is well-known, condition (4.5) also holds for the stan-
dard normal density, the exponential density and more generally the density
of the Gamma distribution, for example. Rb
Since for every density w (x) with support (a, b) , a f (x)2 dx < ∞ implies
p
that f (x) / w (x) ∈ L2 (w) , the following corollary of Theorem 4.1 holds
trivially.

Corollary 4.1. Let L2 (a, b) , −∞ ≤ a < b ≤ ∞, be the Hilbert space of


square integrable Borel measurable real functions on (a, b), with inner product
Rb
hf, gi = a f (x) g (x) dx and associated norm and metric. Every function
f ∈ L2 (a, b) can be written as
̰ !
p X
f (x) = w(x) γk pk (x|w) a.e. on (a, b),
k=0

where w is a density with support (a, b) satisfying the moment conditions


(4.5), the pk (x|w)’s are the orthonormal polynomials generated by w(x) and
56 CHAPTER 4. ORTHOGONAL POLYNOMIALS
p
the γk ’s are the Fourier coefficients of f (x)/ w(x), i.e.,
Z b p
γk = f (x)pk (x|w) w(x)dx.
a

This result implies that the functions


p
ψk (x|w) = pk (x|w) w(x), k ∈ N,

form a complete orthonormal sequence in L2 (a, b):


³n p o∞ ´
2
L (a, b) = span pk (x|w) w(x) .
k=0

Of course, the ψk (x|w)’s are no longer polynomials.


If max (|a| , |b|) < ∞ then there is another way to construct a complete
orthonormal sequence in L2 (a, b) , as follows. Let W (x) be the distribution
function of a density w with bounded support (a, b). Then

G (x) = a + (b − a) W (x)

is a one-to-one mapping of (a, b) onto (a, b), with inverse

G−1 (y) = W −1 ((y − a) / (b − a))

where W −1 is the inverse of W (x) . For every f ∈ L2 (a, b) ,


Z b Z b Z b
2 2
(b − a) f (G (x)) w(x)dx = f (G (x)) dG (x) = f (x)2 dx < ∞.
a a a

Hence f (G (x)) ∈ L2 (w) and thus by Theorem 4.1,


X

f (G (x)) = γk pk (x|w) a.e. on (a, b)
k=0

where
Z b Z b
1
γk = f (G (x)) pk (x|w)w(x)dx = f (G (x)) pk (x|w)dG(x)
a b−a a
Z b
1 ¡ ¢
= f (x) pk G−1 (x) |w dx
b−a a
4.2. THE THREE-TERM RECURRENCE RELATION 57

Consequently

¡ ¡ −1 ¢¢ X

¡ ¢
f (x) = f G G (x) = γk pk G−1 (x) |w a.e. on (a, b)
k=0

Note that dG−1 (x) /dx = dG−1 (x) /dG (G−1 (x)) = 1/G0 (G−1 (x)) , so
that
Z b
¡ ¢ ¡ ¢
pk G−1 (x) |w pm G−1 (x) |w dx
a
Z b
¡ ¢ ¡ ¢ ¡ ¢
= pk G−1 (x) |w pm G−1 (x) |w G0 G−1 (x) dG−1 (x)
a
Z b
= pk (x|w) pm (x|w) G0 (x) dx
a
Z b
= (b − a) pk (x|w) pm (x|w) w (x) dx = (b − a) I (k = m)
a

Thus,

Corollary 4.2. Let w be a density with bounded support (a, b), satisfying
the moment conditions (4.5). Let W be the c.d.f. of w, with inverse W −1 .
Then the functions
¡ ¢ p
ψk (x|w) = pk W −1 ((x − a) / (b − a)) |w / (b − a), k ∈ N0 ,

form a complete orthonormal sequence in L2 (a, b), i.e., every f ∈ L2 (a, b)


P Rb
can be written as f (x) = ∞
k=0 αk ψk (x|w) a.e. on (a, b), where αk = a f (x)
ψk (x|w) dx.

4.2 The three-term recurrence relation


It follows from (4.1) that p0 (x|w) ≡ 1, andRit follows from (4.2) that p1 (x|w) =

α1,0 + x can be constructed by solving −∞ (α1,0 + x) w(x)dx = 0. Hence,
R∞
given that w(x) is a density, α1,0 = − −∞ x.w(x)dx. The question now arises
how to construct these orthogonal polynomials further for k ≥ 2.
The answer is the following.
58 CHAPTER 4. ORTHOGONAL POLYNOMIALS

P
Theorem 4.2. Every sequence of polynomials pk (x|w) = kj=0 αk,j xj , with
αk,k = 1, satisfying the orthogonality condition (4.2), with w(x) satisfying
the moment conditions (4.5), can be generated recursively by the three-term
recurrence relation (hereafter referred to as TTRR)
pk+1 (x|w) + (bk − x) pk (x|w) + ck pk−1 (x|w) = 0, k ∈ N, (4.9)
where R∞
[Link] (x|w)2 w(x)dx
bk = R−∞
∞ (4.10)
−∞
pk (x|w)2 w(x)dx
and R∞
pk (x|w)2 w(x)dx
ck = R ∞−∞ (4.11)
p (x|w)2 w(x)dx
−∞ k−1

qR

Next, let dk = p (x|w)2 w(x)dx, so that pk (x|w) = pk (x|w)/dk is a
−∞ k
sequence of orthonormal polynomials. Substituting pk (x|w) = dk .pk (x|w) in
(4.9), (4.10) and (4.11) yields
dk+1 dk−1
pk+1 (x|w) + (bk − x) pk (x|w) + ck p (x|w) = 0, k ≥ 1,
dk dk k−1
R∞
where bk = −∞ [Link] (x|w)2 w(x)dx and ck = d2k /d2k−1 , hence
dk+1 dk
pk+1 (x|w) + (bk − x) pk (x|w) + p (x|w) = 0, k ≥ 1.
dk dk−1 k−1
Moreover, note that
xpk−1 (x|w) dk [Link]−1 (x|w) dk
lim = lim = ,
|x|→∞ pk (x|w) dk−1 |x|→∞ pk (x|w) dk−1
where the latter equality is due to the normalization αk,k = 1 in Theorem
4.2. Thus:

Theorem 4.3. Every sequence pk (x|w) of orthonormal polynomials with


respect to a density function w(x) satisfying the moment conditions (4.5)
can be generated recursively by the TTRR
ak+1 .pk+1 (x|w) + (bk − x) pk (x|w) + ak .pk−1 (x|w) = 0, k ∈ N. (4.12)
4.3. EXAMPLES OF ORTHONORMAL POLYNOMIALS 59

where ¯ ¯
¯ [Link]−1 (x|w) ¯¯
¯
ak = ¯ lim (4.13)
|x|→∞ p (x|w) ¯
k

and Z ∞
bk = [Link] (x|w)2 w(x)dx. (4.14)
−∞

4.3 Examples of orthonormal polynomials


4.3.1 Hermite polynomials
If w(x) is the density of the standard normal distribution,
¡ ¢ √
wN [0,1] (x) = exp −x2 /2 / 2π,

the orthonormal polynomials involved satisfy the TTRR


√ √
k + 1pk+1 (x|wN [0,1] ) − [Link] (x|wN [0,1] ) + kpk−1 (x|wN [0,1] ) = 0, k ∈ N,

starting from p0 (x|wN [0,1] ) = 1, p1 (x|wN [0,1] ) = x. Thus in this case ak = k
and bk = 0 in (4.12). These polynomials are known as Hermite1 polynomials.
It follows from Theorem 4.1 that the Hermite polynomials span the
Hilbert space L2 (wN [0,1] ), and it follows from Corollary 4.1 that
µnq o∞ ¶
2
L (R) = span wN [0,1] (x)pk (x|wN [0,1] ) .
k=0

Consequently, any density f (x) on R can be represented by


̰ !2
X
f (x) = wN [0,1] (x) γk pk (x|wN [0,1] )
k=0

P∞
where k=0 γk2 = 1.
1
Charles Hermite (1822-1901).
60 CHAPTER 4. ORTHOGONAL POLYNOMIALS

4.3.2 Laguerre polynomials


The standard exponential density function

wExp (x) = I (x ≥ 0) exp (−x)

gives rise to the orthonormal Laguerre2 polynomials, with TTRR

(k + 1) pk+1 (x|wExp ) + (2k + 1 − x) pk (x|wExp ) + [Link]−1 (x|wExp ) = 0,

for k ∈ N, starting from p0 (x|wExp ) = 1, p1 (x|wExp ) = x − 1. Thus in this


case ak = k and bk = 2k + 1.
Since the moment conditions (4.5) hold for wExp (x), it followsRfrom Theo-

rem 4.1 that any Borel measurable real function
P f (x) satisfying 0 exp(−x)

f (x)2 dx < ∞R ∞can be written as f (x) = k=0 γk pk (x|wExp ) a.e. on [0, ∞),
where γk = 0 exp(−x)pk+1 (x|w)f (x)dx.
Again, it follows from Corollary 4.1 that
¡ ¢
L2 (0, ∞) = span {exp(−x/2)pk (x|wExp )}∞ k=0 ,

hence any density f (x) on [0, ∞) can be written as


̰ !2
X X

f (x) = exp(−x) γk pk (x|wExp ) , with γk2 = 1.
k=0 k=0

4.3.3 Legendre polynomials


The uniform density on [−1, 1],
1
wU[−1,1] (x) = I (|x| ≤ 1) ,
2
generates the orthonormal Legendre3 polynomials on [−1, 1], with TTRR
k+1
√ √ pk+1 (x|wU[−1,1] ) − [Link] (x|wU[−1,1] ) (4.15)
2k + 3 2k + 1
k
+√ √ pk−1 (x|wU[−1,1] ) = 0,
2k + 1 2k − 1
2
Edmund Nicolas Laguerre (1834-1886)
3
Adrien-Marie Legendre (1752-1833)
4.3. EXAMPLES OF ORTHONORMAL POLYNOMIALS 61

for k ∈ N, starting from p0 (x|wU[−1,1] ) = 1, p1 (x|wU[−1,1] ) = 3x.
Note that the orthonormal Legendre polynomials pk (x|wU[−1,1] ) satisfy

Z 1
pk (2u − 1|wU[−1,1] )pm (2u − 1|wU[−1,1] )du
0
Z
1 1
= p (2u − 1|wU[−1,1] )pm (2u − 1|wU[−1,1] )d (2u − 1)
2 0 k
Z
1 1
= p (x|wU[−1,1] )pm (x|wU[−1,1] )dx = I (k = m)
2 −1 k

Hence,
pk (u|wU[0,1] ) = pk (2u − 1|wU[−1,1] ), k ∈ N0 ,

is a sequence of orthonormal polynomials w.r.t. the uniform density on [0, 1],

wU[0,1] (u) = I (0 ≤ u ≤ 1)

The pk (u|wU[0,1] )’s are known as the shifted Legendre polynomials, also
called the orthonormal Legendre polynomials on the unit interval [0, 1]. Sub-
stituting x = 2u − 1 and pk (x|wU[−1,1] )) = pk (u|wU[0,1] ) in (4.15) yields the
TTRR

(k + 1) /2
√ √ ρk+1 (u|wU[0,1] ) + (0.5 − u) .ρk (u|wU[0,1] )
2k + 3 2k + 1
k/2
+√ √ ρk−1 (u|wU[0,1] ) = 0, k ∈ N,
2k + 1 2k − 1

starting from ρ0 (u) = 1, ρ1 (u) = 3 (2u − 1) .
Again, it follows from Theorem 4.1 that any P Borel measurable real func-

tion f (x) on [0, 1] can be written as f (x) = γ ρ (u|wU[0,1] ), where
R1 2
¡© k k
k=0 ª∞ ¢
γk = 0 f (x) ρk (u|wU[0,1] )dx, hence L (0, 1) = span ρk (u|wU[0,1] k=0 .
These shifted Legendre polynomials have been used by Bierens (2008)
and Bierens and Carvalho (2007) to model semi-nonparametrically the unob-
served heterogeneity of interval-censored mixed proportional hazard models
and bivariate mixed proportional hazard models, respectively.
62 CHAPTER 4. ORTHOGONAL POLYNOMIALS

4.3.4 Chebyshev polynomials


Chebyshev polynomials on [−1, 1]
Chebyshev polynomials on [−1, 1] are generated by the weight function
1
wC[−1,1] (x) = √ I (|x| < 1) . (4.16)
π 1 − x2
This is a density function on (−1, 1). To see this, let θ = arccos(x), so that
x = cos(θ), and observe that

dx p √
= − sin(θ) = − 1 − cos2 (θ) = − 1 − x2 ,

hence
d arccos(x) −1
=√ (4.17)
dx 1 − x2
Then
Z 1 Z
1 1 1
√ dx = − d arccos(x)
−1 π 1 − x2 π −1
arccos(−1) − arccos(1)
= =1
π
because arccos(−1) = π and arccos(1) = 0. Clearly, the corresponding dis-
tribution function is
arccos(−1) − arccos(x)
WC[−1,1] (x) = , x ∈ [−1, 1].
π
The orthogonal (but not orthonormal) Chebyshev polynomials pk (x|wC[−1,1] )
satisfy the TTRR

pk+1 (x|wC[−1,1] ) − 2xpk (x|wC[−1,1] ) + pk−1 (x|wC[−1,1] ) = 0, k ∈ N, (4.18)

starting from p0 (x|wC[−1,1] ) = 1, p1 (x|wC[−1,1] ) = x, with orthogonality prop-


erties

Z 1 ⎨ 0 if k 6= m,
pk (x|wC[−1,1] )pm (x|wC[−1,1] )
√ dx = 1/2 if k = m > 0,
−1 π 1 − x2 ⎩
1 if k = m = 0.
4.3. EXAMPLES OF ORTHONORMAL POLYNOMIALS 63

An important practical difference with the other polynomials discussed


so far is that Chebyshev polynomials have the closed form4 :

pk (x|wC[−1,1] ) = cos (k. arccos(x)) . (4.19)


To see this, observe from (4.17) and the well-known sine-cosine formulas that
Z 1
cos (k. arccos(x)) cos (m. arccos(x))
√ dx
−1 π 1 − x2
Z
1 1
=− cos (k. arccos(x)) cos (m. arccos(x)) d arccos(x)
π −1
Z
1 π
= cos (k.θ) cos (m.θ) dθ
π 0
Z π Z π
1 1
= cos ((k + m) θ) dθ + cos ((k − m) θ) dθ
2π 0 2π 0
µ ¶
1 sin ((k + m) π) sin ((k − m) π)
= +
2 (k + m) π (k − m) π

⎨ 0 if k 6= m,
= 1/2 if k = m > 0, (4.20)

1 if k = m = 0.

Moreover, the TTRR (4.18) follows from

cos ((k + 1) .θ) − 2 cos(θ) cos (k.θ) + cos ((k − 1) .θ)


= cos (k.θ) cos(θ) − sin (k.θ) sin(θ) − 2 cos(θ) cos (k.θ)
+ cos (k.θ) cos(θ) + sin (k.θ) sin(θ) = 0.

Hence, the functions (4.19) satisfy the TTRR (4.18) and are therefore genuine
polynomials.
In view of (4.20) we can now define the orthonormal Chebyshev polyno-
mials as
½
1 for k = 0,
pk (x|wC[−1,1] ) = √
2 cos (k. arccos(x)) for k ∈ N.
¡ √ ¢
4
Note that arccos(x) = atan −x/ 1 − x2 + 12 .π, where atan(x) is the inverse of the
tangents function tan(θ) = sin(θ)/ cos(θ), θ ∈ (−π/2, π/2) . In most programming lan-
guages the function atan(x) is an intrinsic function. For example, in Visual Basic this
function is the ATN(x) function.
64 CHAPTER 4. ORTHOGONAL POLYNOMIALS

It is trivial to verify that the density (4.16) satisfies the moment condi-
tion (4.5), so that the Chebyshev polynomials form a complete orthonormal
sequence in the Hilbert space L2 (wC[−1,1] ) involved.

Further properties of Chebyshev polynomials


Because pn (x|wC[−1,1] ) is a polynomial of order n in x ∈ [−1, 1], it has at most
n real roots in [−1, 1]. Obviously, these roots are

xn,k = cos (π(k − 0.5)/n) , k = 1, 2, ..., n

Moreover,

Lemma 4.1. For j1 , j2 = 0, 1, 2, ..., n − 1,

X
n
cos (πj1 (k − 0.5)/n) cos (πj2 (k − 0.5)/n)
k=1

X
n ⎨ 0 if j1 6= j2 ,
= pj1 (xn,k |wC[−1,1] )pj2 (xn,k |wC[−1,1] ) = n/2 if j1 = j2 > 0,

k=1 n if j1 = j2 = 0.

Now interpret k in Lemma 4.1 as a time index: k = t = 1, ...., n, and


denote

P0,n (t) ≡ 1, Pj,n (t) = 2 cos(jπ(t − 0.5)/n),
j = 1, 2, ..., n − 1, t = 1, 2, ...., n.

The Pj,n (t)’s are known as Chebyshev time polynomials, which by Lemma
4.1 satisfy

1X
n
Pi,n (t)Pj,n (t) = I(i = j), i, j = 0, 1, 2, ..., n − 1.
n t=1

Consequently, any function g(t) of time t = 1, 2, ..., n can be represented by

X
n−1
1X
n
g(t) = cj,n Pj,n (t), where cj,n = g(k)Pj,n (k).
j=1
n k=1
4.3. EXAMPLES OF ORTHONORMAL POLYNOMIALS 65

In particular, if g(t) is smooth then

X
m
g(t) ≈ cj,n Pj,n (t)
j=1

for modest values of m. This approximation has been used in Bierens (1997)
to test the unit root hypothesis against nonlinear trend stationarity, and in
Bierens and Martins (2010) to test for time varying cointegration.

Shifted Chebyshev polynomials

Substituting x = 2u − 1 for u ∈ [0, 1] in (4.16) yields

2 1
wC[0,1] (u) = q = p . (4.21)
π 1 − (2u − 1)2 π u (1 − u)

with corresponding distribution function

WC[0,1] (u) = 1 − π−1 arccos(2u − 1), (4.22)

and shifted Chebyshev polynomials


½
1
√ for k = 0,
pk (u|wC[0,1] ) = (4.23)
2 cos (k. arccos(2u − 1)) for k ∈ N.

Again, it follows from Corollary 4.1 that the orthonormal sequence


½ p
wC[0,1] (u) for k = 0,
ψk (u) = √ p
2 wC[0,1] (u) cos (k. arccos(2u − 1)) for k ∈ N,

is complete in L2 (0, 1). Thus, every function f ∈ L2 (0, 1) can be written as

X

f (u) = γk ψk (u) (4.24)
k=0

R1
a.e. on (0, 1), where γk = 0
f (u) ψk (u)du.
66 CHAPTER 4. ORTHOGONAL POLYNOMIALS

4.4 Bivariate functions


Let w1 (x) and w2 (y) be densities with supports (a1 , b1 ) and (a2 , b2 ), respec-
tively, where ∞ ≤ ai < bi ≤ ∞, i = 1, 2, satisfying the conditions of Theorem
4.1. Consider the space L2 (w1 × w2 ) of bivariate Borel measurable functions
f (x, y) satisfying
Z b1 Z b2
w1 (x)w2 (y)f (x, y)2 dxdy < ∞, (4.25)
a1 a2

endowed with the inner product


Z b1 Z b2
hf, gi = w1 (x)w2 (y)f (x, y)g(x, y)dxdy
a1 a2
p
and associated norm ||f || = hf, f i and metric ||f − g||. Then for any fixed
y ∈ (a2 , b2 ) for which
Z b1
w1 (x)f (x, y)2 dx < ∞, (4.26)
a1
2
we have f (x, y) ∈ L (w1 ), hence
X

f (x, y) = γk (y)pk (x|w1 ) a.e. on (a1 , b1 ). (4.27)
k=0
R b1 P
where γk (y) = a1 w1 (x)f (x, y)pk (x|w1 )dx and ∞ 2
k=0 γk (y) < ∞.
Note that by the Cauchy-Schwarz inequality and (4.25),
Z b2 Z b2 µZ b1 ¶2
2
w2 (y)γk (y) dy = w2 (y) w1 (x)pk (x|w1 )f (x, y)dx dy
a2 a2 a1
Z b1 Z b1 Z b2
2
≤ w1 (x)pk (x|w1 ) dx w1 (x)w2 (y)f (x, y)2 dxdy
a1 a1 a2
Z b1 Z b2
= w1 (x)w2 (y)f (x, y)2 dxdy < ∞
a1 a2
Rb
where the second equality follows from the fact that a11 w1 (x)pk (x|w1 )2 dx =
1, so that γk (y) ∈ L2 (w2 ). Consequently, for each k ∈ N0 we have
X

γk (y) = γk,m pm (y|w2 ) a.e. on (a2 , b2 ), (4.28)
m=0
4.4. BIVARIATE FUNCTIONS 67

where
Z b2
γk,m = w2 (y)γk (y)pm (y|w2 )dy
a2
Z b1 Z b2
= w1 (x)w2 (y)f (x, y)pk (x|w1 )pm (y|w2 )dxdy (4.29)
a1 a2

P
and ∞ 2
m=0 γk,m < ∞.
Moreover, note that due to (4.25) the restriction (4.26) holds a.e. on
(a2 , b2 ), so that

X
∞ X

f (x, y) = γk,m pk (x|w1 )pm (y|w2 ) a.e. on (a1 , b1 ) × (a2 , b2 ),
k=0 m=0

where thePdouble-array
P∞ γk,m of Fourier coefficients are given by (4.29) and
satisfies ∞ k=0
2 2
m=0 k,m < ∞. Consequently, the space L (w1 × w2 ) is a
γ
Hilbert space.
Recall that in the case

w1 (x) = w2 (x) = exp(−x2 /2)/ 2π = wN [0,1] (x)

the polynomials pk (x|wN [0,1] ) are the Hermite polynomials. Then every den-
sity f (x, y) on R2 can be written as
¡ ¢
exp − 12 (x2 + y 2 )
f (x, y) =

Ã∞ ∞ !2
XX
× γk,m pk (x|wN [0,1] )pm (y|wN [0,1] ) (4.30)
k=0 m=0
X
∞ X

2 2
a.e. on R , where γk,m = 1.
k=0 m=0

This is the approach taken by Gallant and Nychka (1987). They consider
SNP estimation of Heckman’s (1979) sample selection model, where the bi-
variate error distribution of the latent variable equations involved is modeled
semi-nonparametrically via the Hermite expansion (4.30) of the error density.
68 CHAPTER 4. ORTHOGONAL POLYNOMIALS

4.5 Appendix: Proofs


4.5.1 Theorem 4.1
P Rb
Let f n (x) = nk=0 γk pk (x|w), where γk = a pk (x|w)f (x)w(x)dx, and observe
that due to condition (4.5), f n ∈ L2 (w) . Next, observe that
Z bà X
n
!2
° °
°f − f n °2 = f (x) − γk pk (x|w) w(x)dx
a k=0
Z b X
n Z b
2
= f (x) w(x)dx − 2 γk pk (x|w)f (x)w(x)dx
a k=0 a

X
n X
n Z b
+ γk1 γk2 pk1 (x|w)pk2 (x|w)w(x)dx
k1 =0 k1 =0 a
Z b X
n
2
= f (x) w(x)dx − γk2 ≥ 0. (4.31)
a k=0

Pn Rb
Hence k=0 γk2 ≤ a
f (x)2 w(x)dx < ∞ for all n ≥ 0, and thus

X

γk2 < ∞. (4.32)
k=0

The latter implies that f n is a Cauchy sequence in L2 (w) because

max(n,m)
° ° X X

lim °f n − f m °2 = lim γk2 ≤ lim γk2 = 0.
min(n,m)→∞ min(n,m)→∞ m→∞
k=min(n,m)+1 k=m

Therefore, there exists a function f ∈ L2 (w) such that


° °
lim °f n − f ° = 0. (4.33)
n→∞

This limit function f can be written as

X
n
f (x) = γk pk (x|w) + εn (x) (4.34)
k=0
4.5. APPENDIX: PROOFS 69

for all n ∈ N, where Z b


lim εn (x)2 w(x)dx = 0. (4.35)
n→∞ a

Proof of (4.7)
To prove (4.7), it suffices to show that
Z b
¡ ¢
exp (i.t.x) f (x) − f (x) w(x)dx = 0 (4.36)
a

for all t ∈ R, because (4.36) implies that f (x) = f (x) a.e. on (a, b) , due to
the uniqueness of the Fourier transform.5 .
It follows from the definition of γm and f that for m ≤ n,

¯Z b ¯ ¯Z b ¯
¯ ¡ ¢ ¯ ¯ ¯
¯ f (x) − f (x) p (x|w)w(x)dx¯ = ¯ ε (x)p (x|w)w(x)dx¯
¯ m ¯ ¯ n m ¯
a
sa
Z b
≤ εn (x)2 w(x)dx,
a

hence by (4.35),
Z b ¡ ¢
f (x) − f (x) pm (x|w)w(x)dx = 0 (4.37)
a

for all m ∈ N. This result implies, by induction, that


Z b
¡ ¢
f (x) − f (x) xm w(x)dx = 0 for all m ∈ N. (4.38)
a

In
P∞ its turn (4.38) implies, together with the well-known equality exp (i.t.x) =
m
m=0 (i.t.x) /m!, that for t ∈ R and all n ∈ N,
Z b ¡ ¢
exp(i.t.x) f (x) − f (x) w(x)dx
a
Z bXn
(i.t.x)m ¡ ¢
= f (x) − f (x) w(x)dx
a m=0 m!
5
See for example Bierens (1994, Theorem 3.1.1, p.50).
70 CHAPTER 4. ORTHOGONAL POLYNOMIALS

Z bà X ∞
!
(i.t.x)m ¡ ¢
+ f (x) − f (x) w(x)dx
a m=n+1
m!
Z b à !
X∞
(i.t.x)m ¡ ¢
= f (x) − f (x) w(x)dx
a m=n+1
m!

If −∞ < a < b < ∞ then by dominated convergence,


Z b
¡ ¢
exp(i.t.x) f (x) − f (x) w(x)dx
a
Z bà X∞
!
(i.t.x)m ¡ ¢
= lim f (x) − f (x) w(x)dx = 0
a n→∞
m=n+1
m!

If a = −∞ and/or b = ∞ we can find for arbitrary ε > 0 a finite lower bound


a(ε) > a and a finite upper bound b(ε) < b such that
¯Z ¯
¯ a(ε) ¡ ¢ ¯
¯ ¯
¯ exp(i.t.x) f (x) − f (x) w(x)dx¯ < ε/2
¯ a ¯
¯Z b ¯
¯ ¡ ¢ ¯
¯ exp(i.t.x) f (x) − f (x) w(x)dx¯¯ < ε/2
¯
b(ε)

whereas by dominated convergence


Z b(ε) ¡ ¢
exp(i.t.x) f (x) − f (x) w(x)dx
a(ε)
Z b(ε) Ã X∞
!
(i.t.x)m ¡ ¢
= lim f (x) − f (x) w(x)dx = 0
a(ε) n→∞
m=n+1
m!

Since ε > 0 is arbitrary, we therefore have in either case that (4.36) holds. It
therefore follows from (4.34) and (4.35) that
Z bà X
n
!2
lim f (x) − γk pk (x|w) w (x) dx = 0. (4.39)
n→∞ a k=0

This completes the proof of (4.7).


4.5. APPENDIX: PROOFS 71

Proof of (4.8)
To prove that (4.7) implies (4.8), let X be a random drawing from w(x).
Then by Chebyshev’s inequality, (4.39) implies
X
n
f (X) = plim γk pk (X|w) (4.40)
n→∞
k=0

As is well-known6 , convergence in probability is equivalent to almost sure


(a.s.) convergence along a further subsequence of an arbitrary subsequence
of n. Thus it follows from (4.40) that for any subsequence nj in N there
exists a further subsequence njm such that for m → ∞,
njm
X a.s.
γk pk (X|w) → f (X). (4.41)
k=0

For each n there exists an m such that njm−1 ≤ n < njm . Hence, there
exists a further subsequence jn of njm such that for jn−1 ≤ n < jn and
n → ∞,
jn
X a.s.
γk pk (X|w) → f (X). (4.42)
k=0

The latter implies that


⎡Ã !2 ⎤ Ã jn !2
Xjn X
n X
E⎣ γk pk (X|w) − γk pk (X|w) ⎦ = E γk pk (X|w)
k=0 k=0 k=n+1
jn
X
≤ γk2
k=jn−1 +1

so that
⎡Ã !2 ⎤
jn jn
X
∞ X X
n X
∞ X
E ⎣ γk pk (X|w) − γk pk (X|w) ⎦ ≤ γk2
n=1 k=0 k=0 n=1 k=jn−1 +1

X

≤ γk2 < ∞
k=0

6
See for example Bierens (2004, Theorem 6.B.3, p.168).
72 CHAPTER 4. ORTHOGONAL POLYNOMIALS

Then by Chebyshev’s inequality,


"¯ jn ¯ #
X∞ ¯X X
n ¯
¯ ¯
Pr ¯ γk pk (X|w) − γk pk (X|w)¯ > ε < ∞
n=1
¯ ¯
k=0 k=0

for all ε > 0, which by the Borel-Cantelli lemma7 implies that for n → ∞
jn
X X
n
a.s.
γk pk (X|w) − γk pk (X|w) → 0. (4.43)
k=0 k=0
P a.s.
Combining (4.42) and (4.43), it follows that nk=0 γk pk (X|w) → f (X) as
n → ∞, which is equivalent to (4.8) because the support of w(x) was assumed
to be (a, b).

4.5.2 Theorem 4.2


Due to the normalization αk,k = 1 it follows that pk+1 (x|w) − [Link] (x|w) is
a polynomial of order k, which can be written as a linear combination of
p0 (x|w), p1 (x|w),...,pk (x|w):

X
k
pk+1 (x|w) − [Link] (x|w) = δj,k pj (x|w) (4.44)
j=0

for example. Then for m ≤ k,


Z ∞ Z ∞
0 = pk+1 (x|w)pm (x|w)w(x)dx − [Link] (x|w)pm (x|w)w(x)dx
−∞ −∞
Xk Z ∞
− δj,k pj (x|w)pm (x|w)w(x)dx
j=0 −∞
Z ∞ Z ∞
= − [Link] (x|w)pm (x|w)w(x)dx − δm,k pm (x|w)2 w(x)dx
−∞ −∞

so that
R∞
−∞
([Link] (x|w)) pk (x|w)w(x)dx
δm,k = − R∞ , m = 0, 1, 2, ..., k.
p (x|w)2 w(x)dx
−∞ m
7
See for example Bierens (2004, Theorem 2.B.2, p. 168).
4.5. APPENDIX: PROOFS 73

Because [Link] (x|w) is a polynomial of order m + 1, it follows that for


m ≤ k − 2, [Link] (x|w) is orthogonal to pk (x|w), hence δm,k = 0 for m =
0, 1, ...., k − 2. Thus it follows from (4.44) that

pk+1 (x|w) − [Link] (x|w) = δk,k pk (x|w) + δk−1,k pk−1 (x|w)


= −bk pk (x|w) − ck pk−1 (x|w)

where R∞
[Link] (x|w)2 w(x)dx
bk = −δk,k = R−∞

−∞
pk (x|w)2 w(x)dx
and
R∞
−∞
[Link]−1 (x|w).pk (x|w)w(x)dx
ck = −δk−1,k = R∞
p (x|w)2 w(x)dx
−∞ k−1
R∞
pk (x|w)2 w(x)dx
= R ∞−∞
p (x|w)2 w(x)dx
−∞ k−1

The last equality


Pk follows from the fact that [Link]−1 (x|w) can be written as
[Link]−1 (x|w) = m=0 βm,k pm (x|w), where βk,k = 1, so that
Z ∞ X
k Z ∞
[Link]−1 (x|w).pk (x|w)w(x)dx = βm,k pm (x|w)pk (x|w)w(x)dx
−∞ m=0 −∞
Z ∞
= βk,k pk (x|w)2 w(x)dx
Z ∞ −∞
= pk (x|w)2 w(x)dx.
−∞

4.5.3 Lemma 4.1


Using the well-known cosine formulas 2 cos(a) cos(b) = cos(a + b) + cos(a − b)
and cos(a − b) = cos(a) cos(b) + sin(a) sin(b) we can write

X
n
pj1 (xn,k |wC[−1,1] )pj2 (xn,k |wC[−1,1] )
k=1
X
n
= cos (πj1 (k − 0.5)/n) cos (πj2 (k − 0.5)/n)
k=1
74 CHAPTER 4. ORTHOGONAL POLYNOMIALS

1X 1X
n n
= cos (π(j1 + j2 )(k − 0.5)/n) + cos (π(j1 − j2 )(k − 0.5)/n)
2 k=1 2 k=1
1 Xn
= cos (0.5π(j1 + j2 )/n) cos (π(j1 + j2 )k/n)
2 k=1

1 Xn
+ sin (0.5π(j1 + j2 )/n) sin (π(j1 + j2 )k/n)
2 k=1

1 Xn
+ cos (0.5π(j1 − j2 )/n) cos (π(j1 − j2 )k/n)
2 k=1

1 Xn
+ sin (0.5π(j1 − j2 )/n) sin (π(j1 − j2 )k/n)
2 k=1

Moreover, using the well-known De Moivre formula exp(i.a) = cos(a) +


i. sin(a) it follows that

1X
n
cos (π.m.k/n)
2 k=1
X
n X
n
= exp (i.πm.k/n) + exp (−i.πm.k/n)
k=1 k=1
X
n Xn
= (exp (i.πm/n))k + (exp (−i.πm/n))k
k=1 k=1
exp (i.πm) − 1 exp (−i.πm) − 1
= exp (i.πm/n) + exp (−i.πm/n)
exp (i.πm/n) − 1 exp (−i.πm/n) − 1
exp (i.πm/(2n))
= (cos (πm) − 1)
exp (i.πm/(2n)) − exp (−i.πm/(2n))
exp (−i.πm/(2n))
− (cos (πm) − 1)
exp (i.πm/(2n)) − exp (−i.πm/(2n))
= cos (πm) − 1

and similarly for m 6= 0,

1X
n
sin (π.m.k/n)
2 k=1
4.5. APPENDIX: PROOFS 75

1X 1X
n n
= exp (i.πm.k/n) − exp (−i.πm.k/n)
i k=1 i k=1
1X 1X
n n
= (exp (i.πm/n))k − (exp (−i.πm/n))k
i k=1 i k=1
1 exp (i.πm) − 1 1 exp (−i.πm) − 1
= exp (i.πm/n) − exp (−i.πm/n)
i exp (i.πm/n) − 1 i exp (−i.πm/n) − 1
1 exp (i.πm/(2n)) + exp (−i.πm/(2n))
= (cos (πm) − 1)
i exp (i.πm/(2n)) − exp (−i.πm/(2n))
cos (πm/(2n))
=− (cos (πm) − 1)
sin (πm/(2n))

Thus, for j1 6= j2 ,
X
n
pj1 (xn,k |wC[−1,1] )pj2 (xn,k |wC[−1,1] )
k=1
1
= cos (0.5π(j1 + j2 )/n) (cos (π(j1 + j2 )) − 1)
2
1
− cos (π(j1 + j2 )/(2n)) (cos (π(j1 + j2 )) − 1)
2
1
+ cos (0.5π(j1 − j2 )/n) (cos (π(j1 − j2 )) − 1)
2
1
− cos (π(j1 − j2 )/(2n)) (cos (π(j1 − j2 )) − 1)
2
=0

whereas for j1 = j2 = j > 0,


X
n
pj (xn,k |wC[−1,1] )pj (xn,k |wC[−1,1] )
k=1

1 Xn
1 Xn
= cos (π.j/n) cos (2π.j.k/n) + sin (π.j/n) sin (2π.j.k/n)
2 k=1
2 k=1
1 1
+ n= n
2 2
The case j1 = j2 = 0 is trivial.
76 CHAPTER 4. ORTHOGONAL POLYNOMIALS
Chapter 5
Trigonometric series

5.1 Cosine series representation


Note that the distribution function WC[0,1] (u) defined by (4.22) has inverse
−1
WC[0,1] (u) = (1 + cos (π (1 − u))) /2. (5.1)

It is now easy to verify from Corollary 4.2, (5.1) and (4.24) that every function
f ∈ L2 (0, 1) can be written as
X

√ ¡ ¢
f (u) = γ0 + γk 2 cos k. arccos(2Wc−1 (u) − 1)
k=1
X


= γ0 + γk 2 cos (kπ (1 − u))
k=1
X


= γ0 + γk (−1)k 2 cos (kπu)
k=1
X∞

= α0 + αk 2 cos (kπu)
k=1

where
Z µ ¯ ¶
1
1 ¯
k
αk = γk (−1) = (−1) k
f (u) pk ¯
(1 + cos (π (1 − u)))¯ wC[0,1] du
0 2
( R
1
f (u) du if k = 0,
= R1
0 √
0
f (u) 2 cos (kπu) du if k ∈ N.

77
78 CHAPTER 5. TRIGONOMETRIC SERIES

Consequently,

Theorem 5.1. The functions


½
1 if k = 0,
κk (u) = √
2 cos (kπu) if k ∈ N,

form a complete orthonormal sequence in L2 (0, 1). Thus, given a function


f ∈ L2 (0, 1) , let
X
n

fn (u) = α0 + αk 2 cos (kπu)
k=1
R1 P∞
where αk = 0
f (u) κk (u) du. Then k=0 αk2 < ∞ and
Z 1 X

2
lim (f (u) − fn (u)) du = lim αk2 = 0.
n→∞ 0 n→∞
k=n+1

Consequently, similar to Theorem 4.1, f can be written as


X


f (u) = α0 + αk 2 cos (kπu) a.e. on (0, 1). (5.2)
k=1

5.2 Fourier analysis


Consider the following sequence of functions on [−1, 1]:

ϕ0 (x) = 1 (5.3)
√ √
ϕ2k−1 (x) = 2 sin (kπx) , ϕ2k (x) = 2 cos (kπx) , k ∈ N.

These functions are know as the Fourier series on [−1, 1]. It is easy to verify
that these functions are orthonormal with respect to the weight function
w (x) = 12 I (|x| ≤ 1) , i.e.,
Z 1
1
ϕm (x) ϕk (x) dx = I(m = k)
2 −1

It is a classical Fourier analysis result that


5.2. FOURIER ANALYSIS 79

Theorem 5.2. The Fourier series {ϕn }∞ 2


n=0 is complete in L (−1, 1) .

The ”official” proof of this result is long and tedious. See for example
Young (1988). However, using Theorem 5.1 this result can be proved some-
what easier, as follows.
We need to show that for an arbitrary function g ∈ L2 (−1, 1) ,
Z
1 1
lim (g (x) − gn (x))2 dx, (5.4)
n→∞ 2 −1

where
X
n
√ X
n

gn (x) = α0 + αk 2 cos (kπx) + βk 2 sin (kπx) (5.5)
k=1 k=1

with Fourier coefficients


Z
1 1
α0 = g (x) dx
2 −1
Z
1 1√
αk = 2 cos (kπx) g (x) dx
2 −1
Z
1 1√
βk = 2 sin (kπx) g (x) dx.
2 −1
Let x = 2u − 1 for u ∈ [0, 1], and denote

f (u) = g (2u − 1) , fn (u) = gn (2u − 1)

Then it follows from the well-known sine-cosine equalities that


Z 1 Z 1
α0 = g (2u − 1) du = f (u)du
0 0
Z 1√
αk = 2 cos (kπ (2u − 1)) g (2u − 1) du
0
Z 1√
k
= (−) 2 cos (2kπu) f (u)du
0
Z 1√
βk = 2 sin (kπ (2u − 1)) g (2u − 1) du
0
Z 1√
k
= (−) 2 sin (2kπu) f (u)du
0
80 CHAPTER 5. TRIGONOMETRIC SERIES

and

fn (u) = gn (2u − 1)
X n
√ X
n

= α0 + αk 2 cos (kπ (2u − 1)) + βk 2 sin (kπ (2u − 1))
k=1 k=1
Xn
√ X
n

= α0 + αk (−)k 2 cos (2kπu) + βk (−)k 2 sin (2kπu)
k=1 k=1

Thus, (5.4) is true if and only


Z 1
lim (f (u) − fn (u))2 du = 0.
n→∞ 0

Theorem 5.2 follows now from the following result, which will be proved in
the appendix to this chapter.

Theorem 5.3.√The functions ϕ0 (u) = 1, ϕk (u) = 2 sin(2kπu) if k ≥ 1 is
odd, ϕk (u) = 2 cos (2kπu) if k ≥ 2 is even, form a complete orthonormal
sequence in L2 (0, 1) .

Although Theorem 5.1 was used to prove Theorem 5.2, Theorem 5.2 can
also be proved independently. See for example Young (1988). Then Theorem
5.1 becomes a corollary of Theorem 5.2, as follows.
Let f (u) ∈ L2 (0, 1) be arbitrary, and let g (x) = f (|x|) . Then g (x) ∈
L2 (−1, 1), with Fourier coefficients
Z Z 1
1 1
α0 = f (|x|) dx = f (u) du
2 −1 0
Z Z 1√
1 1√
αk = 2 cos (kπx) f (|x|) dx = 2 cos (kπu) f (u) du
2 −1 0
Z
1 1√
βk = 2 sin (kπx) f (|x|) dx = 0
2 −1

Hence it follows from Theorem 5.2 that


Z 1Ã X
n

!2
lim f (u) − α0 − αk 2 cos (kπu) du (5.6)
n→∞ 0 k=1
5.3. SINE SERIES REPRESENTATION 81

Z Ã !2
1 1 X
n

= lim f (|x|) − α0 − αk 2 cos (kπx) dx
2 n→∞ −1 k=1
=0

Similar to the proof of Theorem 4.1 is follows now from (5.6) that
X


f (u) = α0 + αk 2 cos (kπu) a.e. on (0, 1),
k=1
R1 R1√
where α0 = 0 f (u) du and αk = 0 2 cos (kπu) f (u) du for k ≥ 1, which is
just the result in Theorem 5.1.

5.3 Sine series representation


Let f (x) be a square integrable function on [−1, 1] such that f (x) = −f (−x),
with a possible discontinuity at x = 0. Then
Z √ Z 1 √
1 1
βk = f (x) 2 sin (kπx) du = f (u) 2 sin (kπu) du
2 −1 0
Z 1 √
1
0 = f (x) 2 cos (kπx) dx
2 −1
Z
1 1
0 = f (x)dx = 0
2 −1
R1 ¡ P √ ¢2
Hence by Theorem 5.2, limn→∞ 21 −1 f (x) − nk=1 βk 2 sin (kπx) dx = 0,
which by the condition f (x) = −f (−x) implies
Z 1
lim (f (u) − fn (u))2 du = 0,
n→∞ 0

where
X
n

fn (u) = βk 2 sin (kπu)
k=1

Moreover, it is easy to verify that


Z 1√ √
2 sin (kπu) 2 sin (mπu) du = I(k = m).
0
82 CHAPTER 5. TRIGONOMETRIC SERIES

Thus, we have the following corollary of Theorem 5.2.



Theorem 5.4. The sine series { 2 sin (kπu)}∞ k=1 is a complete orthonormal
sequence in L2 (0, 1). Consequently, any function f ∈ L2 (0, 1) can be written
as
X


f (u) = βk 2 sin (kπu) a.e. on (0, 1),
k=1
R1 √
where βk = 0
f (u) 2 sin (kπu) du.

Note however that fn (u) will be a poor approximation of f (u) for u


close to zero or one because fn (0) = fn (1) = 0 whereas f (0) and f (1)
may be nonzero. The reason is that in general limu→u0 limn→∞ fn (u) 6=
limn→∞ limu→u0 fn (u).

5.4 How well does the cosine series fit?


5.4.1 Exact Fourier coefficients
To check how well the cosine series fit, consider the function f (u) = u(4−3u)
on [0, 1]. Note that this is a density function. For this function we can derive
the Fourier coefficients involved analytically, as
Z 1
α0 = f (u)du = 1,
0
Z √ √
1 ¡ ¢
αk = f (u) 2 cos(kπu)du = −2 2(kπ)−2 (−1)k + 1
0

This way of approximating densities directly by a series expansion has been


advocated by Kronmal and Tarter (1968). However, a potential problem
with this approach is that in general there is no guarantee that fn (u) ≥ 0.
In the following figures the function
Pn f (u)√= u(4 − 3u) is compared with
its SNP approximation fn (u) = 1 + k=1 αk 2 cos(kπu) (dotted curve) for
n = 4, 8, 12.
5.4. HOW WELL DOES THE COSINE SERIES FIT? 83

Figure 5.1: f (u) = u(4 − 3u) compared with fn (u) for n = 4

Figure 5.2: f (u) = u(4 − 3u) compared with fn (u) for n = 8


84 CHAPTER 5. TRIGONOMETRIC SERIES

Figure 5.3: f (u) = u(4 − 3u) compared with fn (u) for n = 12

We see that fn (u) approximates f (u) quite well, even for n = 4, ex-
cept
Pnfor the tails
√ of fn (u) in the latter case. The reason is that fn0 (u) =
− k=1 αk kπ 2 sin(kπu), so that fn0 (0) = fn0 (1) = 0. As expected, the tail
fit becomes better for larger truncation orders n.

5.4.2 Bivariate SNP regression


Let (Y, X) ∈ R2 be a pair of absolutely continuous random variables satisfy-
ing
E[Y 2 ] < ∞, E[X 2 ] < ∞. (5.7)
We can always write

E [Y |X] = f (X) = α + βX + X 2 r(X), (5.8)

where α+βX is the linear projection of Y on 1 and X, with residual X 2 r(X).


Moreover, given an absolutely continuous distribution function G(x) with
density g(x) > 0 on R and inverse G−1 (u), u ∈ [0, 1], we can write

r(x) = ϕ(G(x)) (5.9)

where
ϕ(u) = r(G−1 (u)) (5.10)
5.4. HOW WELL DOES THE COSINE SERIES FIT? 85

Now let us assume that


Z 1 Z ∞
2
ϕ(u) du = r(x)2 g(x)dx < ∞ (5.11)
0 −∞

so that ϕ ∈ L2 (0, 1). Then by Theorem 5.1, ϕ has the series expansion
X


ϕ(u) = γ + δk 2 cos(kπu) a.e. on [0, 1],
k=1

where Z Z 1√
1
γ= ϕ(u)du, δk = 2 cos(kπu)ϕ(u)du.
0 0
Consequently,
X


f (X) = E[Y |X] = α + βX + γX 2 + X 2 δk 2 cos(kπG(X)) a.s.
k=1

Next, let
X
n

2 2
fn (X) = α + βX + γX + X δk 2 cos(kπG(X)),
k=1
Pn √
and denote rn (x) = k=1 δk 2 cos(kπG(X)). Since by Theorem 5.1,

lim rn (x) = r(x) a.e.,


n→∞

it follows that
lim fn (X) = f (X) a.s.
n→∞

In principle we could specify f (X) directly as f (X) = ϕ(G(x)), but if f (x)


is linear then we need the full series expansion of ϕ to fit f (x) = α + βX,
whereas in the case (5.8) the linear regression model corresponds to r(x) ≡ 0.
A convenient choice for G is the logistic distribution function

G(x) = (1 + exp(−x))−1 ,

which has density g(x) = G(x)(1−G(x)) and inverse G−1 (u) = ln(u/(1−u)).
Since all the moments of the Logistic distribution are finite, the condition
(5.11) allows r(x) to be a polynomial of any order.
86 CHAPTER 5. TRIGONOMETRIC SERIES

To check how well fn (X) fits, let Y = f (X) + U, where X and U inde-
pendent standard normally distributed, and

f (x) = (|x| − 1/4)3 (I (x > 1/4) − I (x < −1/4)) .

The reason for this choice of f (x) is to check whether the cosine series ex-
pansion is able to capture the horizontal part of f (x) for |x| ≤ 1/4.
The following three figures compare f (x) with the SNP-OLS estimates
b
fn (x) of fn (x) for n = 4, 8, 12 and x ∈ [−2, 2], with G the Logistic distribution
function, on the basis of a random sample of size 500 from (Y, X).

Figure 5.4: f (x) compared with its SNP-OLS estimate fb4 (x) on [−2, 2]
5.4. HOW WELL DOES THE COSINE SERIES FIT? 87

Figure 5.5: f (x) compared with its SNP-OLS estimate fb8 (x) on [−2, 2]

Figure 5.6: f (x) compared with its SNP-OLS estimate fb12 (x) on [−2, 2]
88 CHAPTER 5. TRIGONOMETRIC SERIES

Figure 5.7: Comparison of f (x) with its nonparametric kernel regression


estimate fe(x) on [−2, 2]

As a comparison I have also estimated f (x) by nonparametric kernel re-


gression, similar to Bierens and Pott-Buter (1990), with standard normal ker-
nel and bandwidth constant determined by in-sample leaving-one-out cross-
validation over the interval [0.1, 2]. The result for x ∈ [−2, 2] is displayed in
Figure 5.7.
As to the SNP results, note the slight wiggle of fbn (x) in the flat area
|x| < 1/4, whereas the nonparametric kernel regression estimator fe(x) is
smoother in this area. However, in view of the fact that this flat part of f (x)
has been approximated via a linear combination of cosine functions the SNP
approach works better than I expected.

5.5 Appendix: Proof of Theorem 5.3


The orthonormality of the sequence {ϕn }∞ n=0 is easy to verify. The complete-
ness proof employs the following steps.
R 11. Let C0 [0, 1] be the space of continuous
Step functions f (u) on [0, 1] satis-
2
fying 0 f (u)du = 0, endowed with the L (0, 1) topology, and let C0,1 [0, 1]
be the space of continuously differentiable functions F (u) on [0, 1] satisfying
F (0) = F (1) = 0, also endowed with the L2 (0, 1) topology. Note that the
5.5. APPENDIX: PROOF OF THEOREM 5.3 89
Ru
functions in C0,1 [0, 1] take the form F (u) = 0 f (x) dx with f (u) = F 0 (u).
It will be shown that C0,1 [0, 1] ⊂ span({ϕn }∞ n=0 ) .
Step 2. It will be shown that C0 [0, 1] is the closure of C0,1 [0, 1], hence
C0 [0, 1] ⊂ span({ϕn }∞ n=0 ) . It follows then trivially that the space C[0, 1] of
continuous functions on [0, 1] is contained in span({ϕn }∞ n=0 ) .
Step 3. Finally, it will be shown that every function in L2 (0, 1) can be
written as a limit of a sequence of continuous functions, hence L2 (0, 1) is the
closure of C[0, 1], so that L2 (0, 1) = span({ϕn }∞ n=0 ) .

Proof of Step 1
Let fn (u) Rand f (u) be the same as in Theorem 5.1,
R u except that due to the
1
condition 0 f (u)du = 0, α0 = 0, and let Fn (u) = 0 fn (x) dx. Then
Xn
αk √
Fn (u) = 2 sin (kπu)
k=1

[(n+1)/2] [n/2]
X α2k−1 √ X α2k √
= 2 sin ((2k − 1) πu) + 2 sin (2kπu)
k=1
(2k − 1) π k=1
2kπ

and
Z 1
sup |F (u) − Fn (u)| ≤ |f (x) − fn (x)| dx
0≤u≤1 0
sZ
1
≤ (f (x) − fn (x))2 dx = o (1) (5.12)
0

Next, observe that


Z 1√ √
−2 2
2 sin ((2k − 1) πu) du = = γ0,k
0 (2k − 1) π
Z 1√ √
2 sin ((2k − 1) πu) 2 cos ((2m − 1) πu) du = 0
Z 01 √ √
2 sin ((2k − 1) πu) 2 cos (2mπu) du
0
−2 −2
= +
(2 (k + m) − 1) π (2 (k − m) − 1) π
2 4k − 2
=−
π (2 (k + m) − 1) (2 (k − m) − 1)
90 CHAPTER 5. TRIGONOMETRIC SERIES

2 k − 1/2
=− = γm,k
π (k − 1/2)2 − m2
Hence
√ X


2 sin ((2k − 1) πu) = γ0,k + γm,k 2 cos (2mπu)
m=1

a.e. on [0, 1]. Now let


[(n+1)/2] [n/2]
X α2k−1 X α2k √
Fen (u) = γ0,k + 2 sin (2kπu)
k=1
(2k − 1) π k=1
2kπ
⎛ ⎞
[(n+1)/2]
XN X α2k−1 √
+ ⎝ γm,k ⎠ 2 cos (2mπu)
m=1 k=1
(2k − 1) π

√ [(n+1)/2]
X α2k−1
[n/2]
X α2k √
= −2 2 2 2 + 2 sin (2kπu)
k=1
(2k − 1) π k=1
2kπ
⎛ ⎞
[(n+1)/2]
1 X
N X α2k−1 √
− 2 ⎝ ⎠ 2 cos (2mπu)
π m=1 k=1
(k − 1/2)2 − m2

where N ≥ [(n + 1) /2] . Then

⎛ ⎞2
Z 1³ ´2 X
∞ [(n+1)/2]
X
1 α2k−1
Fen (u) − Fn (u) du = ⎝ ⎠
π4 m 2 − (k − 1/2)2
0 m=N+1 k=1
⎛ ⎞2
[(n+1)/2]
1 X
∞ X |α2k−1 |
≤ ⎝ ⎠
4
π m=N+1 m 2 − (k − 1/2)2
k=1
à P !2
1 X
∞ [(n+1)/2]
|α2k−1 |
k=1
≤ 4
π m=N+1 m − ([n/2])2
2

à P[(n+1)/2] !2
1 X

k=1 |α2k−1 |
=
π 4 m=N+1 (m − [n/2]) (m + [n/2])
à 1 P[(n+1)/2] !2
1 X

[n/2] k=1 |α2k−1 |
≤ 4
4π m=N+1 m − [n/2]
5.5. APPENDIX: PROOF OF THEOREM 5.3 91

à 1
P[(n+1)/2] !2
1 X

[n/2] k=1 |α2k−1 |

4π 4 m − [n/2]
m=[n/2]+1
à !⎛ [(n+1)/2]
⎞2
1 X∞
1 ⎝ 1 X
= |α2k−1 |⎠
4π 4 m=1
m2 [n/2] k=1
à ∞ !
1 X 1 1 X 2

≤ α
4π 4 m=1
m2 [n/2] k=1 2k−1
= O (1/n)) (5.13)

Hence by (5.12) and (5.13),


Z 1³ ´2
lim Fen (u) − F (u) du = 0
n→∞ 0

Since Fen ∈ span({ϕn }∞ ∞


n=0 ) it follows that F ∈ span({ϕn }n=0 ) , hence C0,1 [0, 1] ⊂
span({ϕn }∞n=0 ) .

Proof of Step 2
R u f ∈ C0 [0, 1] , and extend f (x) for x > 1 as
Choose an arbitrary function
f (x) = f (1). Let F (u) = 0 f (x) dx and
Z u+1/n
(F (u + n−1 ) − F (u)) 1
fn (u) = = f (x)dx
n−1 n u

Then by continuity

lim |fn (u) − f (u)| ≤ lim sup |f (x) − f (u)| = 0


n→∞ n→∞ u≤x≤u+1/n

pointwise in u ∈ [0, 1]. Moreover,

sup |fn (u) − f (u)| ≤ 2 sup |f (u)| < ∞


0≤u≤1 0≤u≤1

Therefore if follows by bounded convergence that


Z 1
lim (fn (u) − f (u))2 du = 0
n→∞ 0
92 CHAPTER 5. TRIGONOMETRIC SERIES

Since fn ∈ C0,1 [0, 1] ⊂ span({ϕn }∞ ∞


n=0 ) it follows now that C0 [0, 1] ⊂ span({ϕn }n=0 ).
Because the functions in C [0, 1] differ from the functions in C0 [0, 1] by
constants only, it follows that C [0, 1] ⊂ span({ϕn }∞ n=0 ).

Proof of Step 3
Let B be an arbitrary Borel subset of [0, 1] and let
µ ¶ µ ¶
−1 −1
fn (u) = exp −n inf |x − u| − exp −n inf |x − u| ,
x∈B x∈B\B

where B is the closure of B. This function is continuous on [0, 1]. To see


this, note that for u1 , u2 ∈ [0, 1],

inf |x − u1 | ≤ |u2 − u1 | + inf |x − u2 |


x∈B x∈B
inf |x − u2 | ≤ |u2 − u1 | + inf |x − u1 |
x∈B\B x∈B\B

hence ¯ ¯
¯ ¯
¯ inf |x − u2 | − inf |x − u1 |¯ ≤ |u2 − u1 |
¯x∈B x∈B
¯
and similarly,
¯ ¯
¯ ¯
¯ inf |x − u2 | − inf |x − u1 |¯ ≤ |u2 − u1 |
¯x∈B\B x∈B\B
¯

For u ∈ B, inf x∈B |x−u| = 0 and inf x∈B\B |x−u| > 0, hence limn→∞ fn (u) =
1. For u ∈ B\B, inf x∈B |x−u| = 0 and inf x∈B\B |x−u| = 0, hence fn (u) = 0,
and for u ∈ [0, 1]\B, inf x∈B |x − u| > 0 and inf x∈B\B |x − u| > 0, hence
limn→∞ fn (u) = 0. Thus

lim fn (u) = I (x ∈ B) .
n→∞

Since fn (u) ∈ C[0, 1] ⊂ span({ϕn }∞ n=0 ) it follows now that for arbitrary

Borel sets B, I (x ∈ B) ∈ span({ϕn }n=0 ) and so are all simple functions on
[0, 1]. Because functions are Borel measurable if and only if they are limits
of sequences of simple functions, it follows that L2 (0, 1) = span({ϕn }∞
n=0 ) .
Chapter 6
Density and distribution
functions

6.1 Density functions on the unit interval


It follows from Theorem 5.1 that for any
P∞density function h(u) on [0, 1] there
∞ 2
exists a sequence {αk }k=0 satisfying k=0 αk = 1 such that
à !2
X∞

h (u) = α0 + αk 2 cos (kπu) a.e. on (0, 1). (6.1)
k=1

The square guarantees that h (u) ≥ 0. Gallant and Nychka (1987) proposed
a similar series expansion on the basis of Hermite polynomials.
Note that the αk ’s in (6.1) are no longer unique. For example, we can
always write h (u) = fB (u)2 , where for an arbitrary Borel set B in [0, 1],
p
fB (u) = (I (u ∈ B) − I (u ∈
/ B)) h(u). (6.2)

Then the αk ’s in (6.1) take the form


Z p Z p
αk = h(u)κk (u) du − h(u)κk (u) du
B [0,1\B

In particular, we may choose for α0 any


∙ Z 1 Z 1p ¸
p
α0 ∈ − h(u)du, h(u)du . (6.3)
0 0

93
94 CHAPTER 6. DENSITY AND DISTRIBUTION FUNCTIONS
³ R p i
1
If we choose α0 ∈ 0, 0 h(u)du then we can reparametrize the Fourier
coefficients αk as
1
α0 = p P∞ 2
1 + m=1 δm
δk
αk = p P∞ 2 , k ∈ N,
1 + m=1 δm
P∞ 2
where m=0 δm < ∞. Hence,

Theorem 6.1. For any density function h(u) on [0, 1] there exist possibly
P∞
∞ 2
uncountable many sequences {δm }m=1 satisfying m=0 δm < ∞ such that
¡ P √ ¢2
1+ ∞
k=1 δk 2 cos (kπu)
h(u) = P a.e. on (0, 1). (6.4)
1+ ∞ 2
m=1 δm

In particular, (6.4) holds for all sequences δk of the form


R1 p √
(I (u ∈ B) − I (u ∈
/ B)) h(u) 2 cos (kπu) du
δk = 0 R1 p , (6.5)
0
(I (u ∈ B) − I (u ∈
/ B)) h(u)du

where B is any Borel set in [0, 1] satisfying


Z 1 p
(I (u ∈ B) − I (u ∈ / B)) h(u)du > 0.
0

Moreover, the corresponding SNP densities


¡ √ P ¢2
1 + 2 nk=1 δk cos (kπu)
hn (u) = P (6.6)
1 + nm=1 δm 2

satisfy v
Z u X
1 u ∞ 2
|h (u) − hn (u)| du ≤ t5 δk → 0 (6.7)
0 k=n+1

Furthermore, the corresponding SNP distribution functions have the closed


form expressions

Hn (u) = u
6.2. UNIQUENESS OF THE SERIES REPRESENTATION 95
"
1 √ X n
sin (kπu) X 2 sin (2mπu)
n
+ Pn 2
2 2 δk + δm (6.8)
1+ m=1 δm k=1
kπ m=1
2mπ
#
X
n X
k−1
sin ((k + m) πu) Xn Xk−1
sin ((k − m) πu)
+2 δk δm +2 δk δm ,
k=2 m=1
(k + m) π k=2 m=1
(k − m) π

and satisfy v
u X
u ∞ 2
sup |H(u) − Hn (u)| ≤ t5 δk → 0. (6.9)
0≤u≤1
k=n+1

6.2 Uniqueness of the series representation


R1
The density h(u) in Theorem 6.1 can be written as h(u) = η(u)2 / 0
η(v)2 dv,
where
X


η(u) = 1 + δm 2 cos (mπu) a.e. on (0, 1). (6.10)
m=1

Moreover, recall that in general,


R1 √ p
0
(I(u ∈ B) − I(u ∈ [0, 1]\B)) 2 cos (mπu) h(u)du
δm = R1 p ,
0
(I(u ∈ B) − I(u ∈ [0, 1]\B)) h(u)du
Z 1 p
1
p P∞ 2 = (I(u ∈ B) − I(u ∈ [0, 1]\B)) h(u)du.
1 + m=1 δm 0

R1 p
for some Borel set B satisfying 0 (I(u ∈ B) − I(u ∈ [0, 1]\B)) h(u)du > 0,
hence
v
u
p u X∞
η(u) = (I(u ∈ B) − I(u ∈ [0, 1]\B)) h(u)t1 + 2
δm (6.11)
m=1

Similarly, given this Borel set B and the corresponding


R1 δm ’s, the SNP
2 2
density (6.6) can be written as hn (u) = ηn (u) / 0 ηn (v) dv, where

X
n

ηn (u) = 1 + δm 2 cos (mπu)
m=1
96 CHAPTER 6. DENSITY AND DISTRIBUTION FUNCTIONS
v
u
p u X
n
t
= (I(u ∈ B) − I(u ∈ [0, 1]\B)) hn (u) 1 + 2
δm (6.12)
m=1

Now suppose that h(u) is continuous and positive on (0, 1). Moreover,
let S ⊂ [0, 1] be the set with Lebesgue measure zero on which h(u) =
limn→∞ hn (u) fails to hold. Then for any u0 ∈ (0, 1)\S, limn→∞ hn (u0 ) =
h(u0 ) > 0, hence for sufficient large n, hn (u0 ) > 0. Because obviously
hn (u) and ηn (u) are continuous on (0, 1), for such an n there exists a small
εn (u0 ) > 0 such that hn (u) > 0 for all u ∈ (u0 − εn (u0 ), u0 + εn (u0 )) ∩ (0, 1),
and therefore
ηn (u)
I(u ∈ B) − I(u ∈ [0, 1]\B) = p p P (6.13)
hn (u) 1 + nm=1 δm
2

is continuous on (u0 −εn (u0 ), u0 +εn (u0 ))∩(0, 1). Substituting (6.13) in (6.11)
it follows now that η(u) is continuous on (u0 − εn (u0 ), u0 + εn (u0 )) ∩ (0, 1),
hence by the arbitrariness of u0 ∈ (0, 1)/S, η(u) is continuous on (0, 1).
Next, suppose that η(u) takes positive and negative values on (0, 1). Then
by the continuity of η(u) on (0, 1) there exists a u0 ∈ (0, 1) for which η(u0 ) = 0
and thus h(u0 ) = 0, which however is excluded by the condition that h(u) > 0
on (0, 1). Therefore, either η(u) > 0 for all u ∈ (0, 1) or η(u) R 1 < 0 for all
u ∈ (0, 1). However, the latter is excluded because by (6.10), 0 η(u)du = 1.
Thus, η(u) > 0 on (0, 1), so that by (6.11), I(u ∈ B) − I(u ∈ [0, 1]\B) = 1
on (0, 1).
Consequently,

Theorem 6.2. For every continuous and positive valued density h(u) on
(0, 1) the sequence {δm }∞
m=1 in Theorem 6.1 is unique, with
R1√ p
0
2 cos (mπu) h(u)du
δm = R1p .
0
h(u)du

6.3 General representation


Given a continuous distribution function G(x) with support Ξ ⊂ R, any
distribution function F (x) with support contained in Ξ can be written as
6.3. GENERAL REPRESENTATION 97

F (x) = H(G(x)), where H(u) = F (G−1 (u)) is a distribution function on


[0, 1]. Moreover, if F and G are absolutely continuous with densities f
and g, respectively, then H is absolutely continuous with density h(u), and
f (x) = h(G(x))g(x). Therefore, f (x) can be estimated semiparametrically
by estimating h(u) semiparametrically.
In general, the role of the a priori chosen distribution function G is three-
fold:

1. G specifies the support of the unknown distribution functions F in the


semi-nonparametric model;

2. G maps the parameter space F of candidate distributions for F one-to-


one onto a space H(0, 1) of distribution functions on the unit interval,
which enables us to develop a unified inference approach for a wide
range of semi-nonparametric models;

3. G serves as an initial guess for F (x) = H(G(x)). If the guess is right


then H(u) = u. A related interpretation of G is that it serves as a (non-
Bayesian) ”prior” for F, with the estimate Hb of H playing the role of
correction mechanism which converts the prior G into a ”posterior” Fb
for F on the basis of data evidence. Another related interpretation
is that F = G represents a standard parametric model for which the
semi-nonparametric model is a generalization.

It follows now from Theorem 6.1 and (6.22) that

Theorem 6.3. Given an absolutely continuous distribution function G(x)


on R with density g(x), any density function f (x) with support contained in
the support of g (i.e., {x : f (x) > 0} ⊂ {x : g(x) > 0}) can be written as
¡ √ P ¢2
1+ 2 ∞
k=1 δk cos (kπG (x))
f (x) = g(x) P (6.14)
1+ ∞ m=1 δm
2

a.e. on {x : f (x) > 0}. Moreover, the corresponding SNP densities


¡ √ P ¢2
1 + 2 nk=1 δk cos (kπG (x))
fn (x) = g(x) P (6.15)
1 + nm=1 δm 2
98 CHAPTER 6. DENSITY AND DISTRIBUTION FUNCTIONS

satisfy v
Z u X
∞ u ∞ 2
|f (x) − fn (x)| dx ≤ t5 δk → 0.
−∞ k=n+1

Furthermore, the SNP distribution function Fn (x) = Hn (G(x)) satisfies


v
u X
u ∞ 2
sup |F (x) − Fn (x)| ≤ t5 δk → 0,
x
k=n+1

Rx
where F (x) = −∞
f (z)dz and Hn (u) is defined by (6.8).

6.4 Smoothness
The non-Euclidean parameter of a semi-nonparametric econometric model
often takes the form of a density function f (x). Usually it is assumed that
f (x) has certain smoothness and regularity properties, like boundedness, con-
tinuity and differentiability. Also, usually the semiparametric model involved
requires that the support of f (x) is connected, i.e.,

{x ∈ R : f (x) > 0} = (a, b),

where possibly a = −∞ and/or b = ∞. To impose these conditions, we need


to impose corresponding smoothness and regularity conditions on the density
g(x) of the a priori chosen distribution function G(x) and on the density h(u)
in the transformation f (x) = h(G(x))g(x).
Denoting u = G(x), we can write

f (G−1 (u))
h(u) =
g(G−1 (u))
Given that G is chosen such that f and g have the same support (a, b), it
follows that h(u) must have support (0, 1), i.e.,

h(u) > 0 on (0, 1), (6.16)

and if f and g are continuous on (a, b) then h(u) is continuous on (0, 1).
Moreover, if it is known that f (x) < ∞ for each x ∈ (a, b), then f (x)/g(x) <
∞ for each x ∈ (a, b), hence h(u) < ∞ for each u ∈ (0, 1). Furthermore,
6.4. SMOOTHNESS 99

since g(x) is an initial guess of f (x), it is reasonable to assume that g(x) is


sufficiently close to f (x) to guarantee that

lim f (x)/g(x) < ∞, lim f (x)/g(x) < ∞


x↓a x↑b

These tail conditions, together with the condition that f (x) < ∞ for each
x ∈ (a, b), are equivalent to sup0≤u≤1 h(u) < ∞. A sufficient condition for
the latter is that the δk ’s in (6.4) satisfy

X

|δk | < ∞. (6.17)
k=1

This condition is stronger than necessary for sup0≤u≤1 h(u) < ∞ only, be-
cause:
P∞ √
Theorem 6.4. Condition (6.17) implies that k=1 δk 2 cos(kπu) is uni-
formly continuous on [0, 1], hence the corresponding density function h(u) in
(6.4) is then uniformly continuous on [0, 1].1

Note that Theorems 6.2 and 6.4 imply the following corollary.

Theorem 6.5. Suppose thatPh(u) has support (0, 1). If the δk ’s in (6.4) are
confined to those for which ∞
k=1 |δk | < ∞ then they are unique.

Next, suppose that f and g are continuously differentiable on (a, b). Then
h(u) is continuously differentiable on (0, 1). A sufficient condition for the
latter is that
X∞
k|δk | < ∞. (6.18)
k=1

To see this, pick any u ∈ [0, 1] and let ε 6= 0 be so small that u + ε ∈ [0, 1].
Then by the mean value theorem there exists a sequence λk (u, ε) ∈ [0, 1] such
that
¯ ∞ ¯
¯1 X X∞ ¯
¯ ¯
lim sup ¯ δk (cos(kπ(u + ε)) − cos(kπu)) + π kδk sin(kπu)¯
ε→0 ¯ ε ¯
k=1 k=1

1
Which implies that sup0≤u≤1 h(u) < ∞.
100 CHAPTER 6. DENSITY AND DISTRIBUTION FUNCTIONS

X

≤ π lim sup k|δk |. |sin(kπu) − sin(kπ(u + λk (u, ε)ε))|
ε→0
k=1
X
n X

≤ π lim sup k|δk |. |sin(kπu) − sin(kπ(u + λk (u, ε)ε))| + 2π k|δk |
ε→0
k=1 k=n+1
X

= 2π k|δk | → 0 as n → ∞,
k=n+1

hence
̰ !
d X X

d cos(kπu) X∞
δk cos(kπu) = δk = −π kδk sin(kπu)
du k=1 k=1
du k=1

and thus
¡ √ P √ ¢ ¡P∞ √ ¢
0 2π 1 + 2 ∞
k=1 δk 2 cos(kπu) k=1 kδk 2 sin(kπu)
h (u) = − P .
1+ ∞ 2
i=1 δi

Note that h0 (0) = h0 (1) = 0. Moreover, it follows similar to Theorem 6.4 that
h0 (u) is uniformly continuous on [0, 1].
Along the same lines it can be shown that
P
Theorem 6.6. If for some natural number ` ≥ 1, ∞ `
k=1 k |δk | < ∞, then the
density function h(u) in (6.4) is `-times continuously differentiable on [0, 1].

6.5 Bivariate densities


Similar to (4.30) and (6.1), any bivariate density h(u, v) on [0, 1] × [0, 1] can
be written as
Ã
X∞
√ X∞

h(u, v) = α0,0 + αk,0 2 cos(kπu) + α0,m 2 cos(mπv)
k=1 m=1
!2
X
∞ X

√ √
+ αk,m 2 cos(kπu) 2 cos(mπv)
k=1 m=1
P∞ P∞
a.e. on [0, 1] × [0, 1], where k=0 m=0 αk.m = 1, and similar to (6.4) we can
reparametrize the αk,m ’s such that h(u, v) becomes
1
h(u, v) = P∞ 2
P∞ 2
P∞ P∞ 2
1+ k=1 δk,0 + m=1 δ0,m + k=1 m=1 δk,m
6.5. BIVARIATE DENSITIES 101
Ã
X

√ X


× 1+ δk,0 2 cos(kπu) + δ0,m 2 cos(mπv)
k=1 m=1
!2
X
∞ X

√ √
+ δk,m 2 cos(kπu) 2 cos(mπv)
k=1 m=1

a.e. on [0, 1] × [0, 1], where

X
∞ X
∞ X
∞ X

2 2 2
δk,0 + δ0,m + δk,m < ∞. (6.19)
k=1 m=1 k=1 m=1

Note that the marginal densities of h(u, v) take the form of a weighted
sum of univariate densities. In particular, denoting
Z 1
h1 (u) = h(u, v)dv
0
√ P
1+ 2 ∞ δ cos(kπu)
h1,0 (u) = P∞k,0 2
k=1
1 + k=1 δk,0
¡ P √ ¢2
δ0,m + ∞ k=1 δk,m 2 cos(kπu)
h1,m (u) = P∞ 2
k=0 δk,m

it can be shown that


¡ P ¢ P∞ ¡P∞ 2 ¢
1+ ∞ 2
k=1 δk,0 h1,0 (u) + δ h1,m (u)
h1 (u) = P∞ 2 P∞ 2 m=1 P∞k=0Pk,m
∞ 2
.
1 + k=1 δk,0 + m=1 δ0,m + k=1 m=1 δk,m

Of course, the δk,m ’s can be reparametrized such that h1 (u) takes the form
(6.4).
Similar to (6.6), let
1
hn (u, v) = Pn 2
Pn 2
Pn Pn 2
1+ k=1 δk,0 + m=1 δ 0,m + k=1 m=1 δk,m
Ã
Xn
√ X
n

× 1+ δk,0 2 cos(kπu) + δ0,m 2 cos(mπv)
k=1 m=1
!2
X
n X
n
√ √
+ δk,m 2 cos(kπu) 2 cos(mπv)
k=1 m=1
102 CHAPTER 6. DENSITY AND DISTRIBUTION FUNCTIONS

be the truncated version of h(u, v). It is not hard to verify that


Z 1Z 1
lim |h(u, v) − hn (u, v)| dudv = 0.
n→∞ 0 0

Finally, note that any density f (x, y) with support Ξx × Ξy ⊂ R2 can be


represented by

f (x, y) = gx (x)gy (y)h (Gx (x), Gy (y)) a.e. on Ξx × Ξy ,

with truncated version

fn (x, y) = gx (x)gy (y)hn (Gx (x), Gy (y)) ,

where Gx is an a priori chosen absolutely continuous distribution function


with density gx and support Ξx , and Gy is an a priori chosen absolutely
continuous distribution function with density gy and support Ξy .

6.6 Appendix: Proofs


6.6.1 Theorem 6.1
The result (6.8) follows from the well-known sine-cosine formulas. To prove
(6.7), denote
P∞ √
1+
k=1 δk 2 cos (kπu)
f (u) = p P ,
1+ ∞ δ 2
Pn √m=1 m
1 + k=1 δk 2 cos (kπu)
fn (u) = p P
1 + nm=1 δm 2

It follows from the Cauchy-Schwarz inequality that


Z Z
1 ¯ ¯ 1
¯f (u)2 − fn (u)2 ¯ du = |f (u) − fn (u)| . |f (u) + fn (u)| du
0 0
sZ
1
≤ 2 (f (u) − fn (u))2 du (6.20)
0
6.6. APPENDIX: PROOFS 103

Moreover,
Z 1
(f (u) − fn (u))2 du
0
à !à !2
X n
1 1
= 1+ δk2 p P∞ 2 − p Pn
1 + δ 1 + 2
k=1 m=1 m m=1 δm
P∞
5 X 2
2 ∞
k=n+1 δk
+ P ≤ δ (6.21)
1+ ∞ m=1 δm
2 4 k=n+1 k

as is not hard to verify. The result (6.7) now follows from (6.20) and (6.21).
Finally, (6.9) follows from
v
Z 1 u X
u ∞ 2
sup |H(u) − Hn (u)| ≤ |h (x) − hn (x)| dx ≤ t5 δk → 0. (6.22)
0≤u≤1 0 k=n+1

6.6.2 Theorem 6.4


Pick any u ∈ [0, 1] and let ε 6= 0 be so small that u + ε ∈ [0, 1]. Then
¯∞ ¯
¯X ¯
¯ ¯
lim sup ¯ δk (cos(kπ(u + ε)) − cos(kπu))¯
ε→0 ¯ ¯
k=1
¯∞ ¯
¯X ¯
¯ ¯
= lim sup ¯ δk ((cos(kπε) − 1) cos(kπu) − sin(kπε) sin(kπu))¯
ε→0 ¯ ¯
k=1
X
n X

≤ lim sup |δk | (|1 − cos(kπε)| + | sin(kπε)|) + 3 |δk |
ε→0
k=1 k=n+1
X

=3 |δk | → 0 as n → ∞.
k=n+1
P∞
By the compactness of [0, 1] this result implies that k=1 δk cos(kπu) is uni-
formly continuous on [0, 1], and so is h(u).
104 CHAPTER 6. DENSITY AND DISTRIBUTION FUNCTIONS
Chapter 7

Compactness

As said before, the non-Euclidean parameter of a semi-nonparametric econo-


metric model often takes the form of a density and/or distribution function.
See the next section for an example. Similar to parametric nonlinear esti-
mation, these non-Euclidean parameters need to be confined to a compact
metric space. In this chapter it will be shown how to construct such compact
metric spaces.

7.1 General density and distribution functions


Recall that the results in Theorem 6.1 read more generally as follows. Given
a complete orthonormal sequence {ρk }∞ 2
k=0 in L (0, 1) with ρ0 (u) ≡ 1, for
every density function h(u) on [0, 1] there exist uncountable many sequences
{δm }∞
m=1 satisfying
X

2
δm <∞ (7.1)
m=1

such that P∞ 2
(1 + δ ρ (u))
h(u) = P∞m m2
m=1
a.e. (7.2)
1 + m=1 δm
Moreover, recall that this representation does not require any smoothness
conditions. Thus (7.2) holds if h(u) is merely Borel measurable. Furthermore,
denoting
P 2
(1 + nm=1 δm ρm (u))
hn (u) = P (7.3)
1 + nm=1 δm2

105
106 CHAPTER 7. COMPACTNESS

for n ≥ 1, it follows that


v
Z u X
1 u ∞ 2
|h (u) − hn (u)| du ≤ t5 δk → 0 (7.4)
0 k=n+1

as n → ∞.
The condition (7.1) can be imposed by imposing the restrictions |δk | ≤ δ k ,
P 2
where δ k is an a priori chosen positive sequence such that ∞ k=1 δ k < ∞. For
example, let
c
δk = √ , (7.5)
1 + k ln(k)
P 2
for some constant c > 0. It is easy to verify that then ∞ 2 2
k=1 δ k < c +c / ln(2).
These restrictions on the δk ’s also play a key-role in proving compactness:

Theorem 7.1. Let D (0, 1) be the space of densities of the type (7.2) sub-
ject to the restrictions |δk | ≤ δ k for some a priori chosen positive sequence
P 2
δ k satisfying ∞ 1
k=1 δ k < ∞, endowed with the L metric
Z 1
kh1 − h2 k1 = |h1 (u) − h2 (u)| du.
0

Then D (0, 1) is compact. Consequently, the space


½ Z u ¾
H (0, 1) = H(u) = h(v)dv, h ∈ D (0, 1)
0

endowed with the ”sup” metric


kH1 − H2 ksup = sup |H1 (u) − H2 (u)|
0≤u≤1

is compact as well. Moreover, let Dn (0, 1) be the space of SNP densities of


the type (7.3), with D0 (0, 1) the singleton {h(u) ≡ 1}, subject to the same
restrictions on the δk ’s, and endowed with the same metric as D (0, 1) . Then
the sequence Dn (0, 1) is dense in D (0, 1):
D (0, 1) = ∪∞
n=0 Dn (0, 1).

Consequently, the spaces


½ Z u ¾
Hn (0, 1) = Hn (u) = hn (v)dv, hn ∈ Dn (0, 1)
0
7.1. GENERAL DENSITY AND DISTRIBUTION FUNCTIONS 107

endowed with the sup metric are dense in H (0, 1):

H (0, 1) = ∪∞
n=0 Hn (0, 1).

Bierens (2008, Theorem 8) proved this result for the case where the ρm (u)
are Legendre polynomials and δ k is given by (7.5). However, as will be shown
below these results hold for any complete orthonormal sequence ρn (u) and
P 2
any positive sequence δ k satisfying ∞ k=1 δ k < ∞.
Similarly, it is easy to construct compact metric spaces of general density
and distribution functions on R. In particular, recall that any density f (x)
with support X ⊂ R can be written as f (x) = h(G(x))g(x), where G(x) is
a given absolutely continuous distribution function with density g(x) and
support containing X: X ⊂ {x ∈ R : g(x) > 0} . Thus, denoting

D (G) = {f (x) = h(G(x))g(x) : h ∈ D (0, 1)}


½ Z x ¾
F (G) = F (x) = f (z)dz : f ∈ D (G)
−∞

it follows trivially fromRTheorem 7.1 that D (G) is a compact metric space



of densities with metric −∞ |f1 (x) − f2 (x)| dx, and F (G) is a compact metric
space of absolutely continuous distribution functions with metric supx∈R |F1 (x)
− F2 (x)|. Moreover, denoting

Dn (G) = {f (x) = h(G(x))g(x) : h ∈ Dn (0, 1)}


½ Z x ¾
Fn (G) = F (x) = f (z)dz : f ∈ Dn (G)
−∞

it follows from Theorem 7.1 that D (G) = ∪∞ ∞


n=0 Dn (G) and F (G) = ∪n=0 Fn (G).
The compactness part of Theorem 7.1 follows from the following two
lemmas and the fact that similar to (6.7), for each pair h1 , h2 ∈ D(0, 1) there
exist sequences δ1 = {δ1,k }∞ ∞
k=1 and δ2 = {δ2,k }k=1 such that

v
Z u∞
1 √ uX
|h1 (u) − h2 (u)| du ≤ 5t (δ1,k − δ2,k )2 .
0 k=1
108 CHAPTER 7. COMPACTNESS

Lemma 7.1. Let {δ k }∞ k=1 be an a priori chosen positive sequence satisfying


P∞ 2 ∞
k=1 δ k < ∞, and let ∆ = Xk=1 [−δ k , δ k ]. Endow the space ∆ with the metric
v
u∞
uX
d(δ1 , δ2 ) = t (δ1,k − δ2,k )2 ,
k=1

where δ1 = {δ1,k }∞ ∞
k=1 ∈ ∆, δ2 = {δ2,k }k=1 ∈ ∆. Then ∆ is compact.

Lemma 7.2. Let s(δ1 , δ2 ) be another metric on ∆ such that for some con-
stant c > 0, s(δ1 , δ2 ) ≤ c.d(δ1 , δ2 ). Then under the conditions of Lemma 7.2,
the space ∆ endowed with the metric s is compact as well.

7.2 Smooth densities on the unit interval


P 2 P∞ `
Note that if we replace the condition ∞
k=1 δ k < ∞ in Lemma 7.1 by k=1 k δ k
< ∞ for some integer ` ≥ 0 and the metric d(δ1 , δ2 ) by
X

d(δ1 , δ2 ) = k` |δ1,k − δ2,k |
k=1

then the result of Lemma 7.1 carries over. Consequently, the following results
hold.

Theorem 7.2. Let D` (0, 1) be the space of densities of the type (6.4) sub-
ject to the restrictions
P∞ ` |δk | ≤ δ k for some a priori chosen positive sequence
δ k satisfying k=1 k δ k < ∞ for some integer ` ≥ 0. Endow D` (0, 1) with
the Sobolev 1 metric
¯ ¯
¯ (m) (m) ¯
kh1 − h2 k` = max sup ¯h1 (u) − h2 (u)¯ , (7.6)
0≤m≤` 0≤u≤1

where h(m) (u) = dm h(u)/(du)m for m ≥ 1, h(0) (u) = h(u). Then D` (0, 1) is
compact. Moreover, let D`,n (0, 1) be the space of SNP densities of the type
(6.6), subject to the same restrictions on the δk ’s, and endowed with the same
metric as D` (0, 1) . Again, D`,0 (0, 1) is the singleton {h(u) ≡ 1}. Then the
sequence D`,n (0, 1) is dense in D` (0, 1): D` (0, 1) = ∪∞n=0 D`,n (0, 1).
1
See for example Adams and Fournier (2003).
7.3. APPENDIX: PROOFS 109

This result follows from the fact that for each pair h1 , h2 ∈ D` (0, 1) with
corresponding sequences {δ1,k }∞ ∞
k=1 and {δ2,k }k=1 we have
̰ !
X
kh1 − h2 k` = O k` |δ1,k − δ2,k | , (7.7)
k=1

as is not hard to verify.

7.3 Appendix: Proofs


7.3.1 Lemma 7.1
To prove the compactness of ∆ it suffices to prove that ∆ is complete and
totally bounded. See Royden (1968, Proposition 15, p.164).
Completeness means that every Cauchy sequence in ∆ takes a limit in ∆.
To show this, let δn = {δn,k }∞ k=1 be an arbitrary Cauchy sequence in ∆, i.e.,
v
u∞
uX
lim d(δn , δm ) = lim t (δn,k − δm,k )2 = 0.
min(n,m)→∞ min(n,m)→∞
k=1

Then for each k ≥ 1, limmin(n,m)→∞ |δn,k − δm,k | = 0, hence δn,k is a Cauchy


sequence in [−δ k , δ k ] and therefore takes a limit δk ∈ [−δ k , δ k ]. Consequently,
δ = {δk }∞
k=1 ∈ ∆ and limn→∞ d (δn , δ) = 0, where the latter follows from

X

lim sup (d (δn , δ))2 = lim sup (δn,k − δk )2
n→∞ n→∞
k=1
X
m
2
X

2
≤ lim sup (δn,k − δk ) + 4 δk
n→∞
k=1 k=m+1
X

2
= 4 δ k → 0 as m → ∞.
k=m+1

Thus, ∆ is complete.
To prove
qPtotal boundedness, let ε > 0 be arbitrary, and choose an n so
∞ 2
large that k=n+1 δ k < ε/4. Denote
¡ ¢ ¡ ¢
∆n = Xnk=1 [−δ k , δ k ] × X∞
k=n+1 {0} . (7.8)
110 CHAPTER 7. COMPACTNESS

Since Xnk=1 [−δ k , δ k ] is a closed and bounded subset of Rn it is compact, hence


∆n is compact. Therefore, there exist elements δ1 , ..., δM of ∆n such that
∆n ⊂ ∪M i=1 {δ∗ ∈ ∆n : d(δ, δi ) < q ε/2}. Since for each δ ∈ ∆ there exists
P∞ 2
a δ∗ ∈ ∆n such that d(δ, δ∗ ) ≤ 2 k=n+1 δ k < ε/2, it follows now that
each δ ∈ ∆ belongs to one of the open sets {δ ∈ ∆ : d(δ, δi ) < ε}, hence
∆ ⊂ ∪Mi=1 {δ ∈ ∆ : d(δ, δi ) < ε}. Thus, ∆ is totally bounded.

7.3.2 Lemma 7.2


Let ∆O be a set which is open under the metric s(., .) but not under the metric
d(., .). Let δ ∈ ∆O be a point of closure under the d-metric. Note that by
assumption, δ is an interior point of ∆O under the s-metric. Then for every
ε > 0 there exists a δ ∈/ ∆O such that d(δ, δ) < ε. But then s(δ, δ) < ε/c,
which would imply that δ is a point of closure under the s-metric as well. This
contradiction implies that open sets under the s-metric are also open under
the d-metric. Consequently, any open covering of ∆ under the s-metric is an
open covering under the d-metricy. Since in the latter case ∆ is compact,
there exists a finite sub-covering of ∆, which is also a finite sub-covering
under the s-metric. Hence ∆ is compact under the s-metric.
Part II
Semi-Nonparametric models
(To be done)

111
113

References
Adams, R.A. & J.J.F. Fournier (2003), Sobolev Spaces. Academic Press.
Andrews, D.W.K. (1994), ”Asymptotics for Semiparametric Economet-
ric Models Via Stochastic Equicontinuity”, Econometrica 62, 43-72.
Bickel, P.J., C.A.J. Klaassen, Y. Ritov & J.A. Wellner (1998), Efficient
and Adaptive Estimation for Semiparametric Models, Springer.
Bierens, H. J. (1997), ”Testing the Unit Root with Drift Hypothesis
Against Nonlinear Trend Stationarity, with an Application to the U.S. Price
Level and Interest Rate”, Journal of Econometrics 81, 29-64.
Bierens, H.J. (2004), Introduction to the Mathematical and Statistical
Foundations of Econometrics. Cambridge University Press.
Bierens, H.J. (2008), ”Semi-Nonparametric Interval-Censored Mixed Pro-
portional Hazard Models: Identification and Consistency Results”, Econo-
metric Theory 24, 749-794.
Bierens, H.J. (2011), ”Consistency and Asymptotic Normality of Sieve
Estimators Under Weak and Verifiable Conditions”. [Link]
~hbierens/[Link]
Bierens, H.J. & J. R. Carvalho (2007), ”Semi-Nonparametric Competing
Risks Analysis of Recidivism”, Journal of Applied Econometrics 22, 971-993.
Bierens, H. J. & L. F. Martins (2010), ”Time Varying Cointegration”,
Econometric Theory 26, 1453-1490.
Bierens, H.J. & and H. Pott-Buter (1990), ”Specification of Engel Curves
by Nonparametric Regression (with discussion)”, Econometric Reviews 9,
123-184.
Billingsley, P. (1968), Convergence of Probability Measures. John Wiley.
Chen, X. (2007), ”Large sample sieve estimation of semi-nonparametric
models”. In J.J. Heckman & E. Leamer (eds.), Handbook of Econometrics,
Vol. 6, Ch. 76. Elsevier.
Chen, X., O. Linton & I. Van Keilegom (2003), ”Estimation of Semipara-
metric Models when the Criterion Function is Not Smooth”, Econometrica,
71, 1591-1608.
Eastwood, B.J. & A.R. Gallant (1991), ”Adaptive Rules for Semi-Non-
parametric Estimators that Achieve Asymptotic Normality”, Econometric
Theory 7, 307-340.
Elbers, C. & G. Ridder (1982), ”True and Spurious Duration Depen-
dence: The Identifiability of the Proportional Hazard Model”, Review of
Economic Studies 49, 403-409.
114

Gallant, A. R. (1981), ”On the Bias in Flexible Functional Forms and an


Essentially Unbiased Form: The Fourier Flexible Form”, Journal of Econo-
metrics 15, 211-245.
Gabler, S, F. Laisney & M. Lechner (1993), ”Seminonparametric Esti-
mation of Binary-Choice Models with an Application to Labor-Force Partic-
ipation”, Journal of Business & Economic Statistics 11, 61-80.
Gallant, A.R. & D.W. Nychka (1987), ”Semi-Nonparametric Maximum
Likelihood Estimation”, Econometrica 55, 363-390.
Gill, R.D. (1989), ”Non- and Semi-Parametric Maximum Likelihood Es-
timators and the Von Mises Method (Part 1)”, Scandinavian Journal of Sta-
tistics 16, 97-128.
Grenander, U. (1981), Abstract Inference. Wiley.
Hamming, R.W. (1973), Numerical Methods for Scientists and Engi-
neers. Dover Publications.
Hannan, E.J., & B.G. Quinn (1979), ”The Determination of the Order
of an Autoregression”, Journal of the Royal Statistical Society, Series B, 41,
190-195.
Heckman, J. J. (1979), ”Sample Selection Bias as a Specification Error”,
Econometrica 47, 153-161.
Horowitz, J.L. (1998), Semiparametric Methods in Econometrics, Springer.
Jennrich, R.I. (1969), ”Asymptotic Properties of Nonlinear Least Squares
Estimators”, Annals of Mathematical Statistics 40, 633-643.
Kronmal, R. & M. Tarter (1968), ”The Estimation of Densities and
Cumulatives by Fourier Series Methods”, Journal of the American Statistical
Association 63, 925-952.
Lancaster, T. (1979), ”Econometric Methods for the Duration of Unem-
ployment”, Econometrica 47, 939-956.
Manski, C.F. (1988), ”Identification of Binary Response Models”, Jour-
nal of the American Statistical Association 83, 729-738.
Newey, W.K. (1997), ”Convergence Rates and Asymptotic Normality for
Series Estimators”, Journal of Econometrics 79, 147-168.
Royden, H.L. (1968), Real Analysis. Macmillan.
Schwarz, G. (1978), ”Estimating the Dimension of a Model”, Annals of
Statistics 6, 461-464.
Young, N. (1988), An Introduction to Hilbert Space. Cambridge Univer-
sity Press.
Shen, X. (1997), ”On the Method of Sieves and Penalization”, Annals of
Statistics 25, 2555-2591.
115

Sims, C.A. (1980), ”Macroeconomics and Reality”, Econometrica 48,


1-48.
White, H. & J. Wooldridge (1991), ”Some Results on Sieve Estimation
with Dependent Observations”. In W.A. Barnett, J. Powell & G. Tauchen
(eds.), Non-parametric and Semi-parametric Methods in Econometrics and
Statistics, Ch. 18, Cambridge University Press.
Wold, H. (1938), A Study in the Analysis of Stationary Time Series.
Almqvist and Wiksell, Sweden.

View publication stats

You might also like