0% found this document useful (0 votes)
10 views18 pages

BayesianEstimation

The document outlines the content of a course on Bayesian Parameter Estimation at MIT, covering various topics such as posterior characterization, Bayes risk criteria, and estimation methods. It emphasizes the Bayesian approach to hypothesis testing and parameter estimation, illustrating concepts with examples from fields like computer vision and autonomous vehicles. The course aims to provide a comprehensive understanding of both discrete and continuous-valued parameters within a Bayesian framework.

Uploaded by

temotoloraia2013
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views18 pages

BayesianEstimation

The document outlines the content of a course on Bayesian Parameter Estimation at MIT, covering various topics such as posterior characterization, Bayes risk criteria, and estimation methods. It emphasizes the Bayesian approach to hypothesis testing and parameter estimation, illustrating concepts with examples from fields like computer vision and autonomous vehicles. The course aims to provide a comprehensive understanding of both discrete and continuous-valued parameters within a Bayesian framework.

Uploaded by

temotoloraia2013
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Massachusetts Institute of Technology

Department of Electrical Engineering and Computer Science


6.7800 Inference and Information
Spring 2026

6 Bayesian Parameter Estimation

6.1 Posterior Characterization of Unknown Parameters . . . . . . . . . 97


6.2 Bayes Risk Criteria for Estimation . . . . . . . . . . . . . . . . . . 101
6.3 Minimum Absolute-Error Estimation . . . . . . . . . . . . . . . . . 102
6.4 Maximum A Posteriori Estimation . . . . . . . . . . . . . . . . . . 103
6.5 Bayes Least-Squares Estimation . . . . . . . . . . . . . . . . . . . 104
6.5.1 Performance Characteristics of BLS Estimators . . . . . . . . . . . 106
6.5.2 Orthogonality Characterization of BLS Estimators . . . . . . . . . 107
6.6* Robust Bayesian Estimation . . . . . . . . . . . . . . . . . . . . . . 108
6.7 Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109
6.8 Implementation Architectures and Hypothesis Classes . . . . . . . 110
6.A* Appendix: Covariance Inequalities . . . . . . . . . . . . . . . . . . 111

The problem of hypothesis testing can be viewed as one of making decisions about the
value of a binary variable taking values in H = {H0 , H1 }. In the Bayesian and minimax
formulations, this variable H is modeled as random, while in the Neyman-Pearson
formulation, H is modeled as deterministic, but unknown.
As mentioned in Section 2.4*, hypothesis testing can be extended to M > 2
hypotheses in a fairly straightforward way, though we will not develop the details of
such M -ary hypothesis testing in these notes. Instead we will proceed from a different
perspective. In particular, an M -ary hypothesis test can be viewed as a problem of
estimating a discrete-valued parameter of the model for the observations. Thus, we
now turn our attention to problems of estimation, where the parameter(s) of interest
are either discrete- or continuous-valued.
As with hypothesis testing, there are both Bayesian and nonBayesian approaches.
In this section, we focus on the Bayesian formulation, where the (generally vector-
valued) parameter of interest is modeled as a random variable x taking values in
some alphabet X. As before, our observations y are random and take values in some
alphabet Y.
Such problems arise in an enormously wide range of applications. For example, in
a computer vision setting, x might represent an image, and y might represent a noisy
or otherwise corrupted version of the image. Or x might represent a 3D scene, and y
might be a pair of (stereo) digital images of the scene. As another example, in the
operation of an autonomous vehicle, x might be a vector representing the position,

96
6.1 Posterior Characterization of Unknown Parameters 97

velocity, and acceleration of the vehicle, and y might be a vector of measurements


from several sensors, including GPS receivers.
Our treatment generally encompasses both discrete- and continuous-valued param-
eters. We emphasize the continuous-valued case, but we note in advance that the
development for discrete-valued parameters is almost identical. We will indicate when
there are significant distinctions. Moreover, as in hypothesis testing, our treatment is
agnostic to whether the data are discrete- or continuous-valued. When not otherwise
specified, we will take X and Y to be real-valued alphabets of appropriate dimensions.

6.1 Posterior Characterization of Unknown Parameters


In the Bayesian framework, there is an a priori distribution px (·) for the unknown
parameter. This represents our belief about x prior to any observation of the measure-
ment y. In addition, the data model py|x (·|·) fully specifies the way in which y contains
information about x. From these quantities, we compute the posterior according to
Bayes’ rule, i.e.,
py|x (y|x) px (x)
px|y (x|y) = R ,
py|x (y|x′ ) px (x′ ) dx′

when the variables have PDFs, for example, which corresponds to revising our belief
about x to take into account the data y = y. As such, within this Bayesian framework,
px|y (·|y) is a complete characterization of our knowledge about x given the data.34 A
few simple examples illustrate this belief revision process and provide some initial
intuition.

Example 6.1. Suppose y ∈ {H, T} represents the outcome of a coin toss where the
probability of H is unknown, and which we model as a random variable x ∈ [0, 1], i.e.,

py |x (y|x) = x1y=H (1 − x)1y=T , y ∈ {H, T}, x ∈ [0, 1].

Suppose x ∼ U([0, 1]) is our prior, i.e.,

px (x) = 1, 0 ≤ x ≤ 1.

Then given a particular observation y = y, the posterior distribution for x is


(
x y=H
px|y (x|y) ∝ py |x (y|x) px (x) = x1y=H (1 − x)1−1y=H =
1 − x y = T,

34
Although we did not emphasize this perspective at the time, clearly our analysis of binary
Bayesian hypothesis testing represents a special case of this analysis, where |X| = 2.

Material protected by copyright. Restrictions on usage and reproduction.


98 Bayesian Parameter Estimation

2 2
<latexit sha1_base64="F5rhB7TI1MSu+pqkwqwsU00xZpo=">AAADBnicnVLNTttAEN64f+D+QTn2YjVC6qWxnSLgiNQLR1AJICVWNN6MzYr1rrU7bhRZuVfqtX2N3hDXvkafoq/QtfGBACdGsvbzzHwzs99OWkphKYr+9rwnT589f7G27r989frN243Nd6dWV4bjiGupzXkKFqVQOCJBEs9Lg1CkEs/Syy9N/OwbGiu0OqFFiUkBuRKZ4EDOdTycbvSjQdRacB/EHeizzo6mm71/k5nmVYGKuARrx3FUUlKDIcElLv1JZbEEfgk5jh1UUKBN6nbSZbDtPLMg08Z9ioLWe5tRQ2GbiF2tUwBdcJBJjZXlRpS0GqaUIHdNHLnJdG2aw94tbBdF+lBsXFG2n9RClRWh4jdTZpUMSAeNZsFMGOQkFw6Aa+8uGvALMMDJKetPWmIdjqz7C/P5PITdwc7nvfBraYTKh1HI3WtZ/KQ0oQ0LV0K3edMZZs1FH81vhnh0ga75LRmUFoowNyCdlnOwCyfY0vfdisR3F+I+OB0O4t1BdLzTP9jvlmWNvWcf2EcWsz12wA7ZERsxzpD9YD/ZL++799u78q5vUr1ex9liK+b9+Q9UY/hh</latexit> <latexit sha1_base64="F5rhB7TI1MSu+pqkwqwsU00xZpo=">AAADBnicnVLNTttAEN64f+D+QTn2YjVC6qWxnSLgiNQLR1AJICVWNN6MzYr1rrU7bhRZuVfqtX2N3hDXvkafoq/QtfGBACdGsvbzzHwzs99OWkphKYr+9rwnT589f7G27r989frN243Nd6dWV4bjiGupzXkKFqVQOCJBEs9Lg1CkEs/Syy9N/OwbGiu0OqFFiUkBuRKZ4EDOdTycbvSjQdRacB/EHeizzo6mm71/k5nmVYGKuARrx3FUUlKDIcElLv1JZbEEfgk5jh1UUKBN6nbSZbDtPLMg08Z9ioLWe5tRQ2GbiF2tUwBdcJBJjZXlRpS0GqaUIHdNHLnJdG2aw94tbBdF+lBsXFG2n9RClRWh4jdTZpUMSAeNZsFMGOQkFw6Aa+8uGvALMMDJKetPWmIdjqz7C/P5PITdwc7nvfBraYTKh1HI3WtZ/KQ0oQ0LV0K3edMZZs1FH81vhnh0ga75LRmUFoowNyCdlnOwCyfY0vfdisR3F+I+OB0O4t1BdLzTP9jvlmWNvWcf2EcWsz12wA7ZERsxzpD9YD/ZL++799u78q5vUr1ex9liK+b9+Q9UY/hh</latexit>

px|y (x|H)

px|y (x|T)
<latexit sha1_base64="l85ji+sGzQOQpBrNVn6OanIt9fs=">AAADB3iclVJNb9QwEPWGrxK+tuUGl4gVUjmwSUrV9sChEheORbBtpU0UTZzZrFXHjmynSxRW4sof4VpOiCs/ggP/BifNYbdFQozk+GVm3jzP2GnJmTZB8Hvg3Lh56/adjbvuvfsPHj4abm4da1kpihMquVSnKWjkTODEMMPxtFQIRcrxJD1708ZPzlFpJsUHU5cYF5ALNmMUjHUlwydl0kTq/OMn+6mX25Fu4Rwh0y+S4SgYB51510HYgxHp7SjZHPyKMkmrAoWhHLSehkFp4gaUYZTj0o0qjSXQM8hxaqGAAnXcdE0svefWk3kzqewSxuu8q4wGCt1G9HqdAsycAo8brDRVrDTrYZMayK2IJbeZVqbd9NXCui7Sv8WmlZkdxA0TZWVQ0MtTziruGem14/QyppAaXlsAVt426tE5KKDGDt2NOmLjT7T98/PFwoe98e6rff99qZjIdwKf2ovU+FJIg9ovbAnZ5SUZztpG/8EPVymt7v9weomVZoVkwmCugNuJLUDXdixL17UPIbx67dfB8c443BsH73ZHh6/7J7FBnpJnZJuEZJ8ckrfkiEwIJZ/JV3JBvjlfnAvnu/PjMtUZ9JzHZM2cn38AUCH5Pw==</latexit>

<latexit sha1_base64="BtrAtKVyOLE/ttxfwpcDAjKUZP0=">AAADB3iclVJNb9QwEPWGrxK+tnCDS8QKqRzYJKVqOXCoxIVjEWxbaRNFE+8kteo4kT3pEoWVuPJHuJYT4sqP4MC/wUn3sNsiIUZy/DIzb55n7LSSwlAQ/B44167fuHlr47Z75+69+w+Gmw8PTVlrjhNeylIfp2BQCoUTEiTxuNIIRSrxKD1908WPzlAbUaoP1FQYF5ArkQkOZF3J8HGVtJE++/jJfprFVmQ6SCCkeZ4MR8E46M27CsIlGLGlHSSbg1/RrOR1gYq4BGOmYVBR3IImwSUu3Kg2WAE/hRynFioo0MRt38TCe2Y9My8rtV2KvN67ymihMF3ErNcpgE44yLjF2nAtKloPU0qQWxFL7jKtTLeZy4VNU6R/i01ryl7FrVBVTaj4xSmzWnpUet04vZnQyEk2FoCVt416/AQ0cLJDd6Oe2PoTY//8fD73YXe883LPf19pofLtwOf2Ig2+UCWh8QtbouzzkhlmXaP/4IerlE73fzhLiZVmVSkUYa5B2onNwTR2LAvXtQ8hvHztV8Hh9jjcHQfvdkb7r5dPYoM9YU/ZFgvZHttnb9kBmzDOPrOv7Jx9c744585358dFqjNYch6xNXN+/gGNNflX</latexit>
0 0
0 x
<latexit sha1_base64="4xd7yr85mY12Gefeqwe6Vxx9TRw=">AAADBnicnVK9TtxAEN5zSADnB0hKGosTUpqcbUCQIgVSmpSg5ADpzjqN98ZmxXrX2h3ncrKuj5Q2eY10Udq8Rp4ir5C1ccEBFSNZ+3lmvpnZbyctpbAURX973qOVx09W19b9p8+ev9jY3Hp5ZnVlOA65ltpcpGBRCoVDEiTxojQIRSrxPL1638TPP6OxQqtPNC8xKSBXIhMcyLlOv0w2+9Egai24C+IO9FlnJ5Ot3r/xVPOqQEVcgrWjOCopqcGQ4BIX/riyWAK/ghxHDioo0CZ1O+ki2HWeaZBp4z5FQeu9yaihsE3ELtcpgC45yKTGynIjSloOU0qQuyaO3GS6Ns1hbxe28yK9LzaqKHub1EKVFaHi11NmlQxIB41mwVQY5CTnDoBr7y4a8EswwMkp649bYh0OrfsL89kshMPBwf5R+LE0QuV7Ucjda1l8ozShDQtXQrd5kylmzUUfzG+GeHCBrvkNGZQWijA3IJ2WM7BzJ9jC992KxLcX4i442xvEh4Po9KB//K5bljW2zXbYaxazI3bMPrATNmScIfvGvrMf3lfvp/fL+32d6vU6ziu2ZN6f/w3y+Ks=</latexit>

1 0 x
<latexit sha1_base64="4xd7yr85mY12Gefeqwe6Vxx9TRw=">AAADBnicnVK9TtxAEN5zSADnB0hKGosTUpqcbUCQIgVSmpSg5ADpzjqN98ZmxXrX2h3ncrKuj5Q2eY10Udq8Rp4ir5C1ccEBFSNZ+3lmvpnZbyctpbAURX973qOVx09W19b9p8+ev9jY3Hp5ZnVlOA65ltpcpGBRCoVDEiTxojQIRSrxPL1638TPP6OxQqtPNC8xKSBXIhMcyLlOv0w2+9Egai24C+IO9FlnJ5Ot3r/xVPOqQEVcgrWjOCopqcGQ4BIX/riyWAK/ghxHDioo0CZ1O+ki2HWeaZBp4z5FQeu9yaihsE3ELtcpgC45yKTGynIjSloOU0qQuyaO3GS6Ns1hbxe28yK9LzaqKHub1EKVFaHi11NmlQxIB41mwVQY5CTnDoBr7y4a8EswwMkp649bYh0OrfsL89kshMPBwf5R+LE0QuV7Ucjda1l8ozShDQtXQrd5kylmzUUfzG+GeHCBrvkNGZQWijA3IJ2WM7BzJ9jC992KxLcX4i442xvEh4Po9KB//K5bljW2zXbYaxazI3bMPrATNmScIfvGvrMf3lfvp/fL+32d6vU6ziu2ZN6f/w3y+Ks=</latexit>

1
Figure 6.1: The posterior distributions for probability of heads x when we observe the
outcome y of a single toss of a coin with unknown bias, given a uniform prior, as in
Example 6.1. The left figure depicts the posterior distribution when we observe heads
(H); the right figure depicts the posterior distribution when we observe tails (T).

so, after incorporating the normalization, we have


(
2x y=H
px|y (x|y) =
2(1 − x) y = T.

The result, depicted in Fig. 6.1, is intuitive. Before we observe a coin flip, we consider
any bias equally likely. But after we observe a coin flip we consider some more likely
than others. In particular, if we observe H, our posterior reflects that we consider
larger values of x more likely, while if we observe T, our posterior reflects the opposite.

Example 6.2. Suppose our waiting time for the arrival of a train y is an exponential
random variable with unknown mean 1/x, which we model as a random variable. In
particular,
py |x (y|x) = xe−xy , y > 0.
Moreover, our prior for x is, itself, an exponential random variable, with mean 1/µ,
i.e.,
px (x) = µe−µx , x > 0,
where µ > 0 is given. Then given a particular waiting time y = y, the posterior

Material protected by copyright. Restrictions on usage and reproduction.


6.1 Posterior Characterization of Unknown Parameters 99

1.25
y1 < y 2 < y 3
<latexit sha1_base64="Y9kuANZ7TAeIr2eoEak6Ck8Dv0E=">AAACZXicZVDRTtswFHUztrFsbGVMe+FhYRXSnpqkoIImHpB44REEBSRSRbeuGywc27JvKFGUb9krfBJfsN+YaTOJlitZPjq+5/qeM9KCW4yip5b3ZuXtu/erH/yPn9Y+f2mvf72wqjCUDagSylyNwDLBJRsgR8GutGGQjwS7HN0ePb9f3jFjuZLnWGo2zCGTfMIpoKPS9kaZxsnWQbJVpr3m3knbnagbzSp4DeIGdEhTJ+l663cyVrTImUQqwNrrONI4rMAgp4LVflJYpoHeQsYqnmtlsPb97RfstYMScmaH1cxTHWw7ZhxMlHFHYjBjlyQ42R9WXOoCmaRzxaQQAarg2Wkw5oZRFKUDQA13qwT0BgxQdHksjKpywBtUSth6kb/XYKzb30/mW1fhwDpxmE2nIfS7uzt74Zk2XGa9fpi7T5QN6+p//ssWpeISWWZAOJdTsKUtc9fj0o6Xs30NLnrduN/tn+52Dveb3FfJJvlJfpGY7JFDckxOyIBQUpI/5IE8tv56a9437/u81Ws1mg2yUN6Pf5L0tpQ=</latexit>

1 <latexit sha1_base64="9LorHwkZWji4XgkPNCji73yVf7E=">AAACaXicZVHRTtswFHUDbJCxrWUPm9hLtAqJvTQJoIL2hLSXPYJYAYlU0a3rBgvHtuwbShTyNbzCB/EN/ASmLRItV7J1dHyPfc/xQAtuMYoeG97S8sqHj6tr/qf1z1++Nlsbp1YVhrIeVUKZ8wFYJrhkPeQo2Lk2DPKBYGeDq78v52fXzFiu5H8sNevnkEk+4hTQUWnzu06rxFzf3LqtrLcT62CZ7v5Om+2oE00qeA/iGWiTWR2lrcafZKhokTOJVIC1F3GksV+BQU4Fq/2ksEwDvYKMVTzXymDt+1tv2AsHJeTM9quJrzrYcswwGCnjlsRgwi5IcHTQr7jUBTJJp4pRIQJUwYvbYMgNoyhKB4Aa7kYJ6CUYoOgymbuqygEvUSlh63n+RoOxbn4/mU5dhT3rxGE2HofQ7ezt7ocn2nCZ7XTD3D2ibFhXr3+waFEqLpFlBoRzOQZb2jJ3PS7teDHb9+B0pxN3O93jvfbhwSz3VfKT/CLbJCb75JD8I0ekRyipyR25Jw+NJ6/l/fA2p61eY6b5RubKaz8DqJu6gw==</latexit>

px|y (x|y3 )

0.75
<latexit sha1_base64="B+2w5UfChOzgI28vdchMdKaTFDg=">AAACaXicZVHRTtswFHUzGCxjo90emOAlWoXEXpqkoIL2hMTLHplGAYlU0a3rBgvHtuwbShTyNbzCB/EN/MRM20lruZKto+N77HuOh1pwi1H03PDeray+X1v/4H/c+PR5s9n6cm5VYSjrUyWUuRyCZYJL1keOgl1qwyAfCnYxvDl5Pb+4ZcZyJc+w1GyQQyb5mFNAR6XNLZ1Wibm9u3dbWe8l1sEy7f5Im+2oE00reAviOWiTeZ2mrcbPZKRokTOJVIC1V3GkcVCBQU4Fq/2ksEwDvYGMVTzXymDt+7v/sVcOSsiZHVRTX3Ww65hRMFbGLYnBlF2S4PhoUHGpC2SSzhTjQgSogle3wYgbRlGUDgA13I0S0GswQNFlsnBVlQNeo1LC1ov8nQZj3fx+Mpu6CvvWicNsMgmh1znYPwz/aMNl1u2FuXtE2bCu/v3BskWpuESWGRDO5QRsacvc9bi04+Vs34LzbifudXq/D9rHR/Pc18kO+U72SEwOyTH5RU5Jn1BSkwfySJ4aL17L++Ztz1q9xlzzlSyU1/4Lpq26gg==</latexit>

0.5
px|y (x|y2 )

<latexit sha1_base64="yK/A8/mxP5h2jLyU5ZsS5msF7YI=">AAACaXicZVHRTtswFHUDDJbBaNnDJvYSrUKClyYBVNCekPayRxArIJEqunXdYOHYln1DibJ8zV7hg/gGfmKmLRItV7J1dHyPfc/xQAtuMYqeGt7S8sqH1bWP/qf1jc+bzdbWhVWFoaxHlVDmagCWCS5ZDzkKdqUNg3wg2OXg9tfL+eUdM5Yr+QdLzfo5ZJKPOAV0VNr8qtMqMXf3f91W1ruJdbBM47202Y460aSC9yCegTaZ1WnaavxMhooWOZNIBVh7HUca+xUY5FSw2k8KyzTQW8hYxXOtDNa+v/OGvXZQQs5sv5r4qoMdxwyDkTJuSQwm7IIER8f9iktdIJN0qhgVIkAVvLgNhtwwiqJ0AKjhbpSA3oABii6TuauqHPAGlRK2nufvNRjr5veT6dRV2LNOHGbjcQjdzuHBUXiuDZfZfjfM3SPKhnX1+geLFqXiEllmQDiXY7ClLXPX49KOF7N9Dy72O3G30z07bJ8cz3JfI9/JD7JLYnJETshvckp6hJKa/CMP5LHx7LW8b972tNVrzDRfyFx57f+kv7qB</latexit>

0.25 px|y (x|y1 )

0
0 1 2 3 4
x
<latexit sha1_base64="sJTJkvFPUj/Ss6xMa2x/8n3ye+g=">AAACUXicZVBNTxsxFHy79IMu/YBy7GXVCKmn7C6ggHpC6qVHUBtAIhF6cV4WC69t2W8botX+gl7LL+upP6U3TJJKTXiSpdH4jT0zI6uk5zz/E8Ubz56/eLn5Ktl6/ebtu+2d9+fe1E5QXxhl3OUIPSmpqc+SFV1aR1iNFF2Mbr883l/8IOel0d95ZmlYYanlRArkQJ3dXW938m4+n/QpKJagA8s5vd6JPg/GRtQVaRYKvb8qcsvDBh1LoahNBrUni+IWS2pkZY3jNkn2/mOvAtRYkR82c/9tuheYcToxLhzN6Zxdk/DkeNhIbWsmLRaKSa1SNuljqnQsHQlWswBQOBmspOIGHQoO2VeeairkGzZG+XaVv7PofPCfDBaum6zvgzgrp9MMe93Dg6Psm3VSl/u9rAqfGJ+1zb+u1yNqIzVT6VCFlFP0Mz+rwk5ou1jv9ik43+8WvW7v7LBzcrzsfRM+wEf4BAUcwQl8hVPogwCCn/AL7qPf0d8Y4nixGkdLzS6sTLz1AJrIsZw=</latexit>

Figure 6.2: The (second-order) Erlang posterior distribution in Example 6.2 for the
exponential distribution parameter x given data y . Progressively larger values of y
concentrate more probability at smaller values of x.

distribution for x is

px|y (x|y) ∝ py |x (y|x) px (x) = µxe−(µ+y)x , x > 0, y > 0,

which after normalizing yields

px|y (x|y) = (µ + y)2 xe−(µ+y)x , x > 0, y > 0,

which is a (second-order) Erlang distribution with parameter35 λ = µ + y. As Fig. 6.2


reflects, this posterior distributes more probability toward smaller values of x as y
increases, and as such also behaves as we might expect: if we experience a longer
waiting time y, we consider a smaller value of x more likely, and vice-versa.

Example 6.3. Suppose we observe a Gaussian random variable with unknown mean,

35
A second-order Erlang distributed random variable u has PDF

p(u; λ) = λ2 ue−λu , u ∈ U = R+ .

Material protected by copyright. Restrictions on usage and reproduction.


100 Bayesian Parameter Estimation

1
y >0
<latexit sha1_base64="vgBQy52/MgGSorKOjCeJMW16hqg=">AAACV3icZVDRTtswFHXCBl3YBmWPvIRVSDw1CaBS7QFV2guPTKwUiVTo1nWDVce27Bu6KMpH8Apfxtdspi3SWo5k6ej4HvueM9KCW4zjF8/f+PBxc6vxKdj+/OXrzm5z79qqwlDWp0ooczMCywSXrI8cBbvRhkE+Emwwmv58vR88MGO5kr+x1GyYQyb5hFNAJw3K9OA8PYjvdltxO54jfE+SJWmRJS7vmt6PdKxokTOJVIC1t0mscViBQU4Fq4O0sEwDnULGKp5rZbAOgsP/1FtHJeTMDqt5ijo8dMo4nCjjjsRwrq5ZcNIdVlzqApmkC8ekECGq8DVbOOaGURSlI0ANd6uE9B4MUHQNrDxV5YD3qJSw9ar+R4Oxbv8gXWxdRX3rzFE2m0XQaZ+enEVX2nCZHXei3H2ibFRXb42vR5SKS2SZAeFSzsCWtszdjGs7We/2Pbk+bieddufXaavXXfbeIPvkOzkiCTkjPXJBLkmfUDIlj+SJPHsv3l9/028sRn1v6flGVuA3/wE43rNB</latexit>

0.8

<latexit sha1_base64="ZeZiJkZkAta4BPKRoRFrLC9gz00=">AAACanicZVHLTuMwFHXDPCDzKrBhxCaaDhKzaRJABbFCYsOSEVNAIlV167rBwrEt+6YlCuFr2ML/8A98BO4DaVquZOvo+B77nuOeFtxiFD3XvKUPHz99Xl7xv3z99v1HfXXt3KrcUNamSihz2QPLBJesjRwFu9SGQdYT7KJ3czw+vxgyY7mS/7DQrJNBKvmAU0BHdesbulsmZnh757ai2k7sGNph8adbb0TNaFLBexDPQIPM6rS7WjtM+ormGZNIBVh7FUcaOyUY5FSwyk9yyzTQG0hZyTOtDFa+v/Ufe+WghIzZTjkxVgVbjukHA2XckhhM2AUJDg46JZc6RybpVDHIRYAqGNsN+twwiqJwAKjhbpSAXoMBii6UuavKDPAalRK2mudvNRjr5veT6dRl2LZOHKajUQit5t7ufnimDZfpTivM3CPKhlX59gmLFqXiEllqQDiXI7CFLTLX49KOF7N9D853mnGr2fq71zg6mOW+TDbJL7JNYrJPjsgJOSVtQsk9eSCP5Kn24q15P73NaatXm2nWyVx5v18BMPe7QA==</latexit>

0.6 px|y (x|y)

0.4

<latexit sha1_base64="zTgg4zMfSLorQB9OnAF6MBSSV5s=">AAACXnicZVBNSysxFE3H73l+uxHcDBZBN50ZlSquBDcuFa0KTl+5TdMxmElCcqcfDPM/3Oq/erv3U4xtBVsPBA4n9yT3nLYW3GIU/at4c/MLi0vLK/6f1bX1jc2t7QerckNZgyqhzFMbLBNcsgZyFOxJGwZZW7DH9uvV1/1jjxnLlbzHoWbNDFLJu5wCOumvbhWJ6Q3Kw8T2BketzWpUi0YIfpN4QqpkgpvWVuUi6SiaZ0wiFWDtcxxpbBZgkFPBSj/JLdNAXyFlBc+0Mlj6/sEP9dlRCRmzzWIUpgwOnNIJusq4IzEYqTMW7J43Cy51jkzSsaObiwBV8BUx6HDDKIqhI0ANd6sE9AUMUHRFTD1VZIAvqJSw5bQ+0GCs299PxlsXYcM6c5j2+yHUa6cnZ+GdNlymx/Uwc58oG5bFd/GzEaXiEllqQLiUfbBDO8zcjGs7nu32N3k4rsX1Wv32tHp5Pul9meyRfXJIYnJGLsk1uSENQokhb+SdfFT+e4vemrcxHvUqE88OmYK3+wlTLrY4</latexit>

0.2 px (x)

0
-4 -3 -2 -1 0 1 2 3 4
x<latexit sha1_base64="sJTJkvFPUj/Ss6xMa2x/8n3ye+g=">AAACUXicZVBNTxsxFHy79IMu/YBy7GXVCKmn7C6ggHpC6qVHUBtAIhF6cV4WC69t2W8botX+gl7LL+upP6U3TJJKTXiSpdH4jT0zI6uk5zz/E8Ubz56/eLn5Ktl6/ebtu+2d9+fe1E5QXxhl3OUIPSmpqc+SFV1aR1iNFF2Mbr883l/8IOel0d95ZmlYYanlRArkQJ3dXW938m4+n/QpKJagA8s5vd6JPg/GRtQVaRYKvb8qcsvDBh1LoahNBrUni+IWS2pkZY3jNkn2/mOvAtRYkR82c/9tuheYcToxLhzN6Zxdk/DkeNhIbWsmLRaKSa1SNuljqnQsHQlWswBQOBmspOIGHQoO2VeeairkGzZG+XaVv7PofPCfDBaum6zvgzgrp9MMe93Dg6Psm3VSl/u9rAqfGJ+1zb+u1yNqIzVT6VCFlFP0Mz+rwk5ou1jv9ik43+8WvW7v7LBzcrzsfRM+wEf4BAUcwQl8hVPogwCCn/AL7qPf0d8Y4nixGkdLzS6sTLz1AJrIsZw=</latexit>

Figure 6.3: The prior and posterior distributions for the unknown mean x of a Gaussian
random variable y , using a Gaussian prior as described in Example 6.3. As depicted, for
an observation y > 0, the posterior distribution is a Gaussian distribution with positive
mean, and more concentrated.

which we model as itself a Gaussian random variable. Specifically,

py |x (·|x) = N(x, σ 2 ) and px (·) = N(0, σ02 ),

where σ 2 and σ02 are given. Then after some straightforward analysis, we obtain36

σ02 σ 2 σ02
 
px|y (·|y) = N 2 y,
σ +σ02 σ 2 +σ02

In this case, we see that the prior is symmetric—it equally favors positive and negative
values for the unknown mean x. However, given a positive data value y, the revised
belief favors a positive mean, and with increased confidence in the sense that

σ 2 σ02 2
var[x|y = y] = 2 2 ≤ σ0 = var[x],
σ +σ0

as Fig. 6.3 reflects.

36
The required computations are considerably less cumbersome to carry out using valuable forth-
coming insight.

Material protected by copyright. Restrictions on usage and reproduction.


6.2 Bayes Risk Criteria for Estimation 101

6.2 Bayes Risk Criteria for Estimation


In an application of interest, one must sometimes go further than computing this
belief and actually make a decision about (i.e., guess) the value of x. In such cases,
the tools of estimation theory are required, for which we now develop a Bayesian
methodology.37
We use x̂(y) to denote our estimate of x based on observing the measurement
y = y. Note that what we are estimating is actually an entire (generally vector-valued)
function x̂(·). In particular, for each possible observed value y, the quantity x̂(y)
represents the estimate of the corresponding value of x. This function is called the
estimator.
In general, to find a good estimator for x, we need some measure of goodness of
candidate estimators. In other words, we need a suitable performance criterion with
respect to which we optimize our choice of estimator. Ultimately, this criterion needs
to reflect the role that the parameter plays in the application of interest. From this
perspective, while the posterior distribution is invariant to the semantics associated
with the parameter, the design of any useful estimator must take such semantics into
account.
In the Bayesian formulation, we begin by choosing a (scalar-valued) function
C(a, â) that specifies the cost of estimating an arbitrary vector a as â. In turn, we
choose our estimator x̂(·) to be the function minimizing the expected cost, i.e.,
 
x̂(·) = arg min E C(x, f (y)) , (6.1)
f (·)

where, as in the case of hypothesis testing, the objective function is referred to as the
Bayes risk.
Analogous to the case of hypothesis testing, the expectation in (6.1) is over x
and y jointly, and hence x̂(·) is a function that minimizes the cost averaged over all
possible (x, y) pairs. Moreover, also analogous to the case of Bayesian hypothesis
testing, solving for the optimum function x̂(·) in (6.1) can, again, be accomplished on
a pointwise basis, i.e., for each particular value y that is observed, we find the best
possible choice (in the sense of (6.1)) for the corresponding estimate x̂(y). Indeed, we
rewrite our objective function in (6.1) in the form

 
Z +∞ Z +∞
E C(x, f (y)) = C(x, f (y)) px,y (x, y) dx dy
−∞ −∞
Z +∞ Z +∞ 
= C(x, f (y)) px|y (x|y) dx py (y) dy, (6.2)
−∞ −∞

37
The revised belief (i.e., posterior) is sometimes referred to as a “soft” decision, and an estimate
is correspondingly referred to as a “hard” decision. However, such terminology is not universal.

Material protected by copyright. Restrictions on usage and reproduction.


102 Bayesian Parameter Estimation

and note that since py (y) ≥ 0, we minimize (6.2) if we choose x̂(y) to minimize the
term in brackets for each particular value of y, i.e.,

Z +∞
x̂(y) = arg min C(x, a) px|y (x|y) dx. (6.3)
a −∞

Note that when x is discrete-valued, the integral in (6.3) becomes a summation.


As we would expect—and as (6.3) indicates—the estimate depends on the model
only through the posterior px|y (·|y), consistent with our earlier insight that this belief
summarizes everything we need to know about the x and y to construct the optimal
Bayesian estimators for any given cost criterion.
In the sequel, we focus on some examples of common cost criteria. Initially, we
focus on the estimation of scalar-valued variables x from (generally) vector observations
y, but we also discuss additive cost criteria for vector parameters, in which
N
X
C(x, x̂) = Ci (xi , x̂i ) (6.4)
i=1

for some Ci (·, ·), and with xi and x̂i denoting the ith elements of the vectors x and x̂,
respectively. When (6.4) is satisfied, estimating vector variables x can be handled in a
component-wise manner.

6.3 Minimum Absolute-Error Estimation


One possible choice for the cost function is based on a minimum absolute-error (MAE)
criterion. The cost function of interest in this case is

C(a, â) = |a − â|. (6.5)

Claim 6.1. The MAE estimate is the median of the posterior belief px|y (·|y), i.e.,


x̂MAE (y) = median px|y (·|y) .

Proof. Substituting (6.5) into (6.3) we obtain


Z +∞
x̂MAE (y) = arg min |x − a| px|y (x|y) dx
a
−∞
Z a Z +∞ 
= arg min (a − x)px|y (x|y) dx + (x − a)px|y (x|y) dx .
a −∞ a (6.6)

Material protected by copyright. Restrictions on usage and reproduction.


6.4 Maximum A Posteriori Estimation 103

Differentiating the quantity inside braces in (6.6) with respect to a gives, via Leibnitz’
rule, the condition
Z a Z +∞ 
px|y (x|y) dx − px|y (x|y) dx = 0. (6.7)
−∞ a a=x̂MAE (y)

Rewriting (6.7) we obtain


Z x̂MAE (y) Z +∞
1
px|y (x|y) dx = px|y (x|y) dx = . (6.8)
−∞ x̂MAE (y) 2

From (6.8) we see that the x̂MAE (y) is the threshold in x of the posterior belief px|y (x|y)
for which half the probability is located above the threshold and, hence, half is also
below the threshold. Hence, the MAE estimator for x given y = y is the median of
the posterior belief.

A natural generalization of the absolute error cost (6.5) to the case of vector
parameters is
C(a, â) = ∥a − â∥1 , (6.9)
where ∥·∥1 is the L1 norm, defined via
X
∥b∥1 = |bi |,
i

with bi denoting the ith element of b. This is an example of additive cost function,
i.e., satisfies (6.4), and is commonly referred to as the L1 loss.

6.4 Maximum A Posteriori Estimation

As an alternative to that considered in the previous section, consider the minimum


uniform cost (MUC) criterion, whereby
(
1 |a − â| > ϵ
C(a, â) = , (6.10)
0 otherwise

which uniformly penalizes all estimation errors with magnitude bigger than ϵ.

Claim 6.2. In the limit as ϵ → 0, the MUC estimate is the mode of the posterior

Material protected by copyright. Restrictions on usage and reproduction.


104 Bayesian Parameter Estimation

belief px|y (·|y), i.e., it is the maximum a posteriori (MAP) estimate38


x̂MAP (y) = mode px|y (·|y) = arg max px|y (a|y). (6.11)
a

From this claim, we see that the MAP estimator can be viewed as resulting from a
Bayes cost formulation in which all errors are, in an appropriate sense, equally bad.
Proof. Substituting (6.10) into (6.3) we obtain that the MUC estimator satisfies
 Z a+ϵ 
ϵ
x̂MUC (y) = arg min 1 − px|y (x|y) dx
a a−ϵ
Z a+ϵ
= arg max px|y (x|y) dx. (6.12)
a a−ϵ

Note that via (6.12) we see that x̂ϵMUC (y) corresponds to the value of a that makes
P(|x − x̂ϵMUC (y)| < ϵ | y = y) as large as possible. This means finding the interval of
length 2ϵ where the posterior density px|y (x|y) is most concentrated.
If we carry this perspective a little further, we see that if we let ϵ get sufficiently
small then the x̂ϵMUC (y) approaches the point corresponding to the peak of the posterior
density, i.e., the MAP estimate (6.11):

lim x̂ϵMUC (y) = arg max px|y (a|y) = x̂MAP (y). (6.13)
ϵ→0 a

The generalization of the limiting minimum uniform cost to the case of vector
parameters corresponds to a cost function that is commonly referred to as the L0
loss.39

6.5 Bayes Least-Squares Estimation


Perhaps the most popular Bayesian estimator is based on a quadratic cost criterion,
which we now develop. Specifically, we consider the mean-square error (MSE) cost
criterion40
N
2 T
X
C(a, â) = a − â = (a − â) (a − â) = (ai − âi )2 , (6.14)
i=1
38
Note that when the posterior distribution does not have a unique mode, then any maximizer is
an MAP estimate.
39
The corresponding L0 norm of a vector is the number of nonzero elements in it. Despite the
nomenclature, it actually doesn’t satisfy the properties required of norm, but is nevertheless useful.
40
In (6.14), ∥·∥ refers to ∥·∥2 , the (standard Euclidean) L2 norm, and thus this criterion is also
referred to as the (squared) L2 loss.

Material protected by copyright. Restrictions on usage and reproduction.


6.5 Bayes Least-Squares Estimation 105

from which we obtain what is termed the Bayes least-squares (BLS) estimator. Since
this estimator minimizes the mean-square estimation error, it is often alternatively
referred to as the minimum mean-square error (MMSE) estimator and denoted using
x̂MMSE (·).

Claim 6.3. The BLS (or MMSE) estimate is the mean of the posterior belief px|y (·|y),
i.e.,
x̂BLS (y) = E[x|y = y].

Proof. First, substituting (6.14) into (6.3) yields


Z +∞
x̂BLS (y) = arg min (x − a)T (x − a)px|y (x|y) dx. (6.15)
a −∞

We begin with the simpler case of scalar estimation, for which (6.15) becomes
Z +∞
x̂BLS (y) = arg min (x − a)2 px|y (x|y) dx. (6.16)
a −∞

As we did in the case of MAE estimation, we can perform the minimization in


(6.16) by differentiating with respect to a and setting the result to zero to find the
local extrema.41 Differentiating the integral in (6.16) we obtain
Z +∞  Z +∞
∂ 2 ∂
(x − a) px|y (x|y) dx = (x − a)2 px|y (x|y) dx
∂a −∞ −∞ ∂a
Z +∞
= −2 (x − a)px|y (x|y) dx. (6.17)
−∞

Setting (6.17) to zero at a = x̂BLS (y) we see that


Z +∞ 
(x − a)px|y (x|y) dx
−∞ a=x̂BLS (y)
Z +∞ Z +∞
= x px|y (x|y) dx − x̂BLS (y) px|y (x|y) dx
−∞ −∞
Z +∞
= E[x|y = y] − x̂BLS (y) px|y (x|y) dx
−∞
= E[x|y = y] − x̂BLS (y) = 0. (6.18)
41
Indeed the objective function is convex:

∂2
Z +∞  Z +∞
2
(x − a) px|y (x|y) dx = 2 px|y (x|y) dx = 2 > 0.
∂a2 −∞ −∞

Material protected by copyright. Restrictions on usage and reproduction.


106 Bayesian Parameter Estimation

Hence,
x̂BLS (y) = E[x|y = y], (6.19)
i.e., that the BLS or MMSE estimate of x given y = y is the mean of the posterior
belief px|y (·|y).
When x is a vector, it suffices to note that since the cost criterion (6.14) is additive
[cf.(6.4)], the minimum is achieved by minimizing the mean-square estimation error in
each scalar component. Hence, we obtain

x̂BLS (y) = E[x|y], (6.20)

from which we see that the BLS estimate of x given y = y is in general the mean of
the posterior belief px|y (·|y).

6.5.1 Performance Characteristics of BLS Estimators

The error
e(x, y) = x̂(y) − x, (6.21)
of any estimator x̂(y) can be expressed in the form

e(x, y) = b + e(x, y) − b , (6.22)

where Z +∞ Z +∞
  
b = E e(x, y) = x̂(y) − x px,y (x, y) dx dy (6.23)
−∞ −∞

is referred to as the (global) bias in the estimator, and where the term in brackets has
zero-mean.
Moreover, the mean-square
 T estimation
 error in such an estimate is the trace of the
error correlation matrix E ee , which can be expressed via (6.23) in the form

E eeT = Λe + bbT ,
 
(6.24)

where   T 
Λe = E e(x, y) − b e(x, y) − b (6.25)
is the error covariance matrix associated with the estimator. Hence we see that in
general, both the bias and covariance contribute to the error correlation and, in turn,
mean-square estimation error.
For the case of a Bayes least-squares estimate specifically, it is straightforward to
verify that there is no bias component to the error, as the following claim asserts.

Material protected by copyright. Restrictions on usage and reproduction.


6.5 Bayes Least-Squares Estimation 107

Claim 6.4. The BLS estimate is unbiased, i.e.,


 
E x̂BLS (y) − x = 0.

Proof. Using (6.23) we have


     
bBLS = E e(x, y) = E x̂BLS (y) − x = E E[x|y] − E[x] = 0, (6.26)

where the last equality follows from a simple application of the law of iterated
expectation.
That the BLS estimate is unbiased means that the MSE performance of the
estimator is simply the trace of the error covariance matrix ΛBLS . Hence, we have the
following performance characterization.
Claim 6.5. The error covariance matrix is the expected covariance of the posterior
belief, i.e.,42
 
ΛBLS = E Λx|y (y) . (6.27)

Proof. Using (6.21), (6.25), and (6.26) we obtain that the associated error covariance
is given by T 
ΛBLS ≜ Λe = E eeT = E x − E[x|y] x − E[x|y] ,
   
(6.28)
where we emphasize that the notation ΛBLS is used to refer to the error covariance
of the BLS estimator. Applying iterated expectation to (6.28) we see that the error
covariance can be written as
h    T i
ΛBLS = E E x − E[x|y] x − E[x|y] y . (6.29)

However, the inner expectation in (6.29) is simply the covariance of the posterior
belief, i.e., Λx|y , whence (6.27).

6.5.2 Orthogonality Characterization of BLS Estimators


Bayes least-squares estimates are unique in having an important orthogonality property.
Specifically, we have the following theorem.
Theorem 6.1 (BLS Orthogonality). An estimator x̂(·) is the Bayes least-squares
estimator, i.e., x̂(·) = x̂BLS (·), if and only if the associated estimation error e(x, y) =
x̂(y) − x is orthogonal to any (vector-valued) function g(·) of the data, i.e.,

E x̂(y) − x g(y)T = 0.
  
(6.30)
42
Note that given an observed value of y, this posterior covariance Λx|y=y is in general a function
of y. As such, we’ll sometimes use the alternative notation Λx|y (y) for this covariance.

Material protected by copyright. Restrictions on usage and reproduction.


108 Bayesian Parameter Estimation

Proof. It is convenient to first rewrite (6.30) as

E x g(y)T = E x̂(y) g(y)T ,


   
(6.31)

and note, using the law of iterated expectation, that the left-hand side of (6.31) can
in turn be expressed in the form
h  i
E x g(y) = E E x g(y) y = E E[x|y] g(y)T .
T T
   
(6.32)

To prove the “only if” statement, choose x̂(·) = x̂BLS (·) in (6.31), and note that in
this case the right-hand expressions in both (6.32) and (6.31) are identical, verifying
(6.30).
To prove the converse, let us rewrite (6.30) using (6.31) and (6.32) as

0 = E x g(y)T − E x̂(y) g(y)T


   

= E E[x|y] g(y)T − E x̂(y) g(y)T


   

= E E[x|y] − x̂(y) g(y)T .


  
(6.33)

Then, since (6.33) must hold for all g(·), let us choose g(y) = E[x|y] − x̂(y) where x̂(·)
is our estimator. In this case (6.33) becomes
  T 
E E[x|y] − x̂(y) E[x|y] − x̂(y) = 0,

from which we can immediately conclude that x̂(y) = E[x|y].43

It is worth emphasizing that Theorem 6.1 ensures what we would expect of an


estimator that yields the minimum mean-square error: that since the error e(x, y) =
x̂(y) − x is uncorrelated with any function of the data we might construct, there is no
further processing that can be done on the data to further reduce the error covariance
in the estimate.
Applications of such orthogonality lead to additional useful matrix inequalities,
examples of which are provided in Appendix 6.A*.

6.6* Robust Bayesian Estimation


While computationally attractive, the quadratic cost function means that the BLS
estimate x̂BLS (y) can be quite sensitive to outliers—data y for which py|x (y|x) is
comparatively small when x is the true value of the parameter. In such cases, a
compromise between the quadratic and absolute error cost criteria is sometimes
43
Here we are using a straightforward consequence of the Chebyshev inequality—that if E zzT = 0
 

then z = 0, or more precisely, P(z = 0) = 1.

Material protected by copyright. Restrictions on usage and reproduction.


6.7 Examples 109

sought, since the latter is much less sensitive to such outliers. One such compromise
is represented by the Huber loss, for which
(
(x − x̂)2 /2 |x − x̂| ≤ δ
Cδ (x, x̂) ≜ (6.34)
δ(|x − x̂| − δ/2) |x − x̂| > δ,

and which is continuously differentiable in |x − x̂|. An even smoother variant is the


pseudo-Huber loss s 
2
(x − x̂)
Cδ′ (x, x̂) ≜ δ 2  1 + − 1 ,
δ2

which matches (6.34) as |x − x̂| → 0 and as |x − x̂| → ∞; in particular, it behaves


like (x − x̂)2 /2 in the former limit, and δ|x − x̂| in the latter one. Yet another variant
is the so-called log-cosh loss

C ′′ (x, x̂) ≜ log cosh(x − x̂) ,




which is twice differentiable everywhere and behaves like (x − x̂)2 /2 as |x − x̂| → 0,


and |x − x̂| − log(2) as |x − x̂| → ∞.
These cost criteria are instances of classical approaches in the field of robust
estimation.

6.7 Examples

Example 6.4 (Revisiting Example 6.1). In this coin toss example, we have, using a
little geometry to calculate the median,
( √
1/ 2 y=H
x̂MAE (y) = √
1 − 1/ 2 y = T.

And clearly (
1 y=H
x̂MAP (y) =
0 y = T.
Finally,
(
Z 1
1y=H 1y=T 2/3 y = H
x̂BLS (y) = E[x|y = y] = x (2x) (2 − 2x) dx =
0 1/3 y = T.

Material protected by copyright. Restrictions on usage and reproduction.


110 Bayesian Parameter Estimation

Example 6.5 (Revisiting Example 6.2). In the train waiting time example, we have

1
x̂MAP = , (6.35)
µ+y

since the mode of a second-order Erlang distribution with parameter λ is 1/λ, while

2
x̂BLS (y) =
µ+y

since the mean of a second-order Erlang distribution with parameter λ is 2/λ. Finally,
we have
c
x̂MAE (y) = ,
µ+y
where the constant c is an involved calculation (and well beyond our scope of interest).
Indeed, the median of an Erlang distribution is c/λ, where
 
1
c = −W−1 − − 1, (6.36)
2e

with W−1 (·) denoting the so-called negative branch of the Lambert W function. Usefully,
it can be shown44 that
5
1.67 ≈ < c ≤ 1 + ln(2) ≈ 1.69,
3
from which we conclude that x̂MAP (y) < x̂MAE (y) < x̂BLS (y).

Example 6.6 (Revisiting Example 6.3). In the Gaussian example, we have

σ02
x̂MAE (y) = x̂MAP (y) = x̂BLS (y) = y, (6.37)
σ 2 +σ02

due to the fact that the Gaussian distribution is both symmetric and unimodal.

6.8 Implementation Architectures and Hypothesis Classes


Even after choosing a cost function and determined the corresponding Bayesian
estimator, its implementation can be a challenge, particularly when x and y are of
very high dimension, as arises in many contemporary applications. For example, in
44
See, e.g., K. P. Choi, “On the medians of Gamma distributions and an equation of Ramanujan,”
Proc. Amer. Math. Soc., vol. 121, no. 1, pp. 245–251, 1994.

Material protected by copyright. Restrictions on usage and reproduction.


6.A* Appendix: Covariance Inequalities 111

a computer vision application, these vectors may have millions of dimensions. As a


result, some flexible architectures are used for the implementation, which generally
cannot represent the optimal estimator exactly but can approximate it. Such an
architecture can be viewed as a function fθ (·)
 whose parameters
 θ∈O - we choose to
optimize the (now constrained) Bayes risk E C x, fθ (y) . In particular, we generate
the estimator
x̂(y) = fθ̂ (y),
where  
θ̂ = arg min E C x, fθ (y) .
θ∈O
-

The family {fθ (·), θ ∈ O}


- is often referred to as the hypothesis class, particularly in
the machine learning community.
In practice, there are many well-established architectures (i.e., functions fθ (·)) from
which to choose for an application of interest. Widely used examples include linear
architectures, whereby

fθ (y) = Ay + d, with θ = (A, d),

as are developed in Section 7, and neural networks, which often use layering to combine
linear processing with so-called nonlinear activation functions in a multi-layer structure.
In modern deep neural networks, θ can include billions of parameters.
In conjunction with choosing an architecture, one also chooses a cost function
C(·, ·). Popular choices include those described earlier, among others. Still others we
will develop axiomatically in subsequent sections.

6.A* Appendix: Covariance Inequalities


BLS estimator analysis leads naturally to useful matrix inequalities, as the following
result establishes.
Theorem 6.2. Let Λe denote the error covariance of an arbitrary estimator x̂(·).
Then the error covariance of the BLS estimator, i.e., ΛBLS , satisfies45

ΛBLS ⪯ Λe , (6.38)

with equality if and only if


 
x̂(y) − E x̂(y) − x = x̂BLS (y) = E[x|y]. (6.39)
45
Recall that for a symmetric matrix A, the notation A ⪰ 0 means that uT Au ≥ 0 for every u,
in which case we say A is positive semidefinite. Moreover, for a pair of symmetric matrices A and
B, the notations A ⪰ B and B ⪯ A are equivalent and interpreted to mean that A − B ⪰ 0. The
notation for positive definiteness is analogous.

Material protected by copyright. Restrictions on usage and reproduction.


112 Bayesian Parameter Estimation

In essence, this states that Bayes least-squares estimator is guaranteed to yield


less uncertainty in the value of x (as measured by the covariance) than any other
estimator—biased or unbiased.
Before proving this result, note that the following corollary is obtained by choosing
x̂(y) = µx (for which Λe = Λx ), and by applying Claim 6.5, and is particularly useful
in applications.
Corollary 6.1.
 
E Λx|y (y) ⪯ Λx (6.40)

with equality if and only if


E[x|y] = E[x]. (6.41)
Note that (6.41) is weaker than the requirement that x and y be independent, but
stronger than the requirement that they be uncorrelated. Note too that it is not true
that Λx|y (y) ⪯ Λx for each y ∈ Y in general. In essence, not every y ∈ Y need be
helpful in estimating x, as the following example illustrates.

Example 6.7. With X = Y = {0, 1}, suppose

px|y (x|0) = 1x=0 , px|y (x|1) = 1/2, x ∈ X, and py (y) = 1/2, y ∈ Y.

Then since a B(p) random variable has variance p(1 − p) it follows that

λx|y (0) = 0 and λx|y (1) = 1/4,

and in turn
1 1 1 1
λx|y = py (0) λx|y (0) + py (1) λx|y (1) = ·0+ · = .
2 2 4 8
Moreover,
(
3/4 x = 0
px (x) = py (0) px|y (x|0) + py (1) px|y (x|1) =
1/4 x = 1,

whence
3 1 3
λx = · = .
4 4 16
As a result, we have the ordering

λx|y (0) < λx|y < λx < λx|y (1) .


| {z } |{z} |{z} | {z }
=0 =1/8 =3/16 =1/4

Material protected by copyright. Restrictions on usage and reproduction.


6.A* Appendix: Covariance Inequalities 113

Finally, we return to the proof.

Proof of Theorem 6.2. To establish (6.38), let

g(y) ≜ x̂(y) − x̂BLS (y) − b, (6.42)

where b denotes the bias in x̂(·), and note that


h  T i
Λe = E x̂(y) − x − b x̂(y) − x − b
h  T i
= E g(y) + x̂BLS (y) − x g(y) + x̂BLS (y) − x
T 
= E g(y) g(y)T + E x̂BLS (y) − x x̂BLS (y) − x
   
T
+ E x̂BLS (y) − x g(y)T + E x̂BLS (y) − x g(y)T .
    
(6.43)

From Theorem 6.1 we get that the last two terms in (6.43) are zero. Using this
together with the definition of ΛBLS we get

Λe − ΛBLS = E g(y) g(y)T .


 
(6.44)

The right-hand side of (6.44) is in general positive semidefinite, which verifies (6.38),
and equal to zero if and only if g(y) = 0, which using (6.42) yields (6.39).

Material protected by copyright. Restrictions on usage and reproduction.

You might also like