BayesianEstimation
BayesianEstimation
The problem of hypothesis testing can be viewed as one of making decisions about the
value of a binary variable taking values in H = {H0 , H1 }. In the Bayesian and minimax
formulations, this variable H is modeled as random, while in the Neyman-Pearson
formulation, H is modeled as deterministic, but unknown.
As mentioned in Section 2.4*, hypothesis testing can be extended to M > 2
hypotheses in a fairly straightforward way, though we will not develop the details of
such M -ary hypothesis testing in these notes. Instead we will proceed from a different
perspective. In particular, an M -ary hypothesis test can be viewed as a problem of
estimating a discrete-valued parameter of the model for the observations. Thus, we
now turn our attention to problems of estimation, where the parameter(s) of interest
are either discrete- or continuous-valued.
As with hypothesis testing, there are both Bayesian and nonBayesian approaches.
In this section, we focus on the Bayesian formulation, where the (generally vector-
valued) parameter of interest is modeled as a random variable x taking values in
some alphabet X. As before, our observations y are random and take values in some
alphabet Y.
Such problems arise in an enormously wide range of applications. For example, in
a computer vision setting, x might represent an image, and y might represent a noisy
or otherwise corrupted version of the image. Or x might represent a 3D scene, and y
might be a pair of (stereo) digital images of the scene. As another example, in the
operation of an autonomous vehicle, x might be a vector representing the position,
96
6.1 Posterior Characterization of Unknown Parameters 97
when the variables have PDFs, for example, which corresponds to revising our belief
about x to take into account the data y = y. As such, within this Bayesian framework,
px|y (·|y) is a complete characterization of our knowledge about x given the data.34 A
few simple examples illustrate this belief revision process and provide some initial
intuition.
Example 6.1. Suppose y ∈ {H, T} represents the outcome of a coin toss where the
probability of H is unknown, and which we model as a random variable x ∈ [0, 1], i.e.,
px (x) = 1, 0 ≤ x ≤ 1.
34
Although we did not emphasize this perspective at the time, clearly our analysis of binary
Bayesian hypothesis testing represents a special case of this analysis, where |X| = 2.
2 2
<latexit sha1_base64="F5rhB7TI1MSu+pqkwqwsU00xZpo=">AAADBnicnVLNTttAEN64f+D+QTn2YjVC6qWxnSLgiNQLR1AJICVWNN6MzYr1rrU7bhRZuVfqtX2N3hDXvkafoq/QtfGBACdGsvbzzHwzs99OWkphKYr+9rwnT589f7G27r989frN243Nd6dWV4bjiGupzXkKFqVQOCJBEs9Lg1CkEs/Syy9N/OwbGiu0OqFFiUkBuRKZ4EDOdTycbvSjQdRacB/EHeizzo6mm71/k5nmVYGKuARrx3FUUlKDIcElLv1JZbEEfgk5jh1UUKBN6nbSZbDtPLMg08Z9ioLWe5tRQ2GbiF2tUwBdcJBJjZXlRpS0GqaUIHdNHLnJdG2aw94tbBdF+lBsXFG2n9RClRWh4jdTZpUMSAeNZsFMGOQkFw6Aa+8uGvALMMDJKetPWmIdjqz7C/P5PITdwc7nvfBraYTKh1HI3WtZ/KQ0oQ0LV0K3edMZZs1FH81vhnh0ga75LRmUFoowNyCdlnOwCyfY0vfdisR3F+I+OB0O4t1BdLzTP9jvlmWNvWcf2EcWsz12wA7ZERsxzpD9YD/ZL++799u78q5vUr1ex9liK+b9+Q9UY/hh</latexit> <latexit sha1_base64="F5rhB7TI1MSu+pqkwqwsU00xZpo=">AAADBnicnVLNTttAEN64f+D+QTn2YjVC6qWxnSLgiNQLR1AJICVWNN6MzYr1rrU7bhRZuVfqtX2N3hDXvkafoq/QtfGBACdGsvbzzHwzs99OWkphKYr+9rwnT589f7G27r989frN243Nd6dWV4bjiGupzXkKFqVQOCJBEs9Lg1CkEs/Syy9N/OwbGiu0OqFFiUkBuRKZ4EDOdTycbvSjQdRacB/EHeizzo6mm71/k5nmVYGKuARrx3FUUlKDIcElLv1JZbEEfgk5jh1UUKBN6nbSZbDtPLMg08Z9ioLWe5tRQ2GbiF2tUwBdcJBJjZXlRpS0GqaUIHdNHLnJdG2aw94tbBdF+lBsXFG2n9RClRWh4jdTZpUMSAeNZsFMGOQkFw6Aa+8uGvALMMDJKetPWmIdjqz7C/P5PITdwc7nvfBraYTKh1HI3WtZ/KQ0oQ0LV0K3edMZZs1FH81vhnh0ga75LRmUFoowNyCdlnOwCyfY0vfdisR3F+I+OB0O4t1BdLzTP9jvlmWNvWcf2EcWsz12wA7ZERsxzpD9YD/ZL++799u78q5vUr1ex9liK+b9+Q9UY/hh</latexit>
px|y (x|H)
px|y (x|T)
<latexit sha1_base64="l85ji+sGzQOQpBrNVn6OanIt9fs=">AAADB3iclVJNb9QwEPWGrxK+tuUGl4gVUjmwSUrV9sChEheORbBtpU0UTZzZrFXHjmynSxRW4sof4VpOiCs/ggP/BifNYbdFQozk+GVm3jzP2GnJmTZB8Hvg3Lh56/adjbvuvfsPHj4abm4da1kpihMquVSnKWjkTODEMMPxtFQIRcrxJD1708ZPzlFpJsUHU5cYF5ALNmMUjHUlwydl0kTq/OMn+6mX25Fu4Rwh0y+S4SgYB51510HYgxHp7SjZHPyKMkmrAoWhHLSehkFp4gaUYZTj0o0qjSXQM8hxaqGAAnXcdE0svefWk3kzqewSxuu8q4wGCt1G9HqdAsycAo8brDRVrDTrYZMayK2IJbeZVqbd9NXCui7Sv8WmlZkdxA0TZWVQ0MtTziruGem14/QyppAaXlsAVt426tE5KKDGDt2NOmLjT7T98/PFwoe98e6rff99qZjIdwKf2ovU+FJIg9ovbAnZ5SUZztpG/8EPVymt7v9weomVZoVkwmCugNuJLUDXdixL17UPIbx67dfB8c443BsH73ZHh6/7J7FBnpJnZJuEZJ8ckrfkiEwIJZ/JV3JBvjlfnAvnu/PjMtUZ9JzHZM2cn38AUCH5Pw==</latexit>
<latexit sha1_base64="BtrAtKVyOLE/ttxfwpcDAjKUZP0=">AAADB3iclVJNb9QwEPWGrxK+tnCDS8QKqRzYJKVqOXCoxIVjEWxbaRNFE+8kteo4kT3pEoWVuPJHuJYT4sqP4MC/wUn3sNsiIUZy/DIzb55n7LSSwlAQ/B44167fuHlr47Z75+69+w+Gmw8PTVlrjhNeylIfp2BQCoUTEiTxuNIIRSrxKD1908WPzlAbUaoP1FQYF5ArkQkOZF3J8HGVtJE++/jJfprFVmQ6SCCkeZ4MR8E46M27CsIlGLGlHSSbg1/RrOR1gYq4BGOmYVBR3IImwSUu3Kg2WAE/hRynFioo0MRt38TCe2Y9My8rtV2KvN67ymihMF3ErNcpgE44yLjF2nAtKloPU0qQWxFL7jKtTLeZy4VNU6R/i01ryl7FrVBVTaj4xSmzWnpUet04vZnQyEk2FoCVt416/AQ0cLJDd6Oe2PoTY//8fD73YXe883LPf19pofLtwOf2Ig2+UCWh8QtbouzzkhlmXaP/4IerlE73fzhLiZVmVSkUYa5B2onNwTR2LAvXtQ8hvHztV8Hh9jjcHQfvdkb7r5dPYoM9YU/ZFgvZHttnb9kBmzDOPrOv7Jx9c744585358dFqjNYch6xNXN+/gGNNflX</latexit>
0 0
0 x
<latexit sha1_base64="4xd7yr85mY12Gefeqwe6Vxx9TRw=">AAADBnicnVK9TtxAEN5zSADnB0hKGosTUpqcbUCQIgVSmpSg5ADpzjqN98ZmxXrX2h3ncrKuj5Q2eY10Udq8Rp4ir5C1ccEBFSNZ+3lmvpnZbyctpbAURX973qOVx09W19b9p8+ev9jY3Hp5ZnVlOA65ltpcpGBRCoVDEiTxojQIRSrxPL1638TPP6OxQqtPNC8xKSBXIhMcyLlOv0w2+9Egai24C+IO9FlnJ5Ot3r/xVPOqQEVcgrWjOCopqcGQ4BIX/riyWAK/ghxHDioo0CZ1O+ki2HWeaZBp4z5FQeu9yaihsE3ELtcpgC45yKTGynIjSloOU0qQuyaO3GS6Ns1hbxe28yK9LzaqKHub1EKVFaHi11NmlQxIB41mwVQY5CTnDoBr7y4a8EswwMkp649bYh0OrfsL89kshMPBwf5R+LE0QuV7Ucjda1l8ozShDQtXQrd5kylmzUUfzG+GeHCBrvkNGZQWijA3IJ2WM7BzJ9jC992KxLcX4i442xvEh4Po9KB//K5bljW2zXbYaxazI3bMPrATNmScIfvGvrMf3lfvp/fL+32d6vU6ziu2ZN6f/w3y+Ks=</latexit>
1 0 x
<latexit sha1_base64="4xd7yr85mY12Gefeqwe6Vxx9TRw=">AAADBnicnVK9TtxAEN5zSADnB0hKGosTUpqcbUCQIgVSmpSg5ADpzjqN98ZmxXrX2h3ncrKuj5Q2eY10Udq8Rp4ir5C1ccEBFSNZ+3lmvpnZbyctpbAURX973qOVx09W19b9p8+ev9jY3Hp5ZnVlOA65ltpcpGBRCoVDEiTxojQIRSrxPL1638TPP6OxQqtPNC8xKSBXIhMcyLlOv0w2+9Egai24C+IO9FlnJ5Ot3r/xVPOqQEVcgrWjOCopqcGQ4BIX/riyWAK/ghxHDioo0CZ1O+ki2HWeaZBp4z5FQeu9yaihsE3ELtcpgC45yKTGynIjSloOU0qQuyaO3GS6Ns1hbxe28yK9LzaqKHub1EKVFaHi11NmlQxIB41mwVQY5CTnDoBr7y4a8EswwMkp649bYh0OrfsL89kshMPBwf5R+LE0QuV7Ucjda1l8ozShDQtXQrd5kylmzUUfzG+GeHCBrvkNGZQWijA3IJ2WM7BzJ9jC992KxLcX4i442xvEh4Po9KB//K5bljW2zXbYaxazI3bMPrATNmScIfvGvrMf3lfvp/fL+32d6vU6ziu2ZN6f/w3y+Ks=</latexit>
1
Figure 6.1: The posterior distributions for probability of heads x when we observe the
outcome y of a single toss of a coin with unknown bias, given a uniform prior, as in
Example 6.1. The left figure depicts the posterior distribution when we observe heads
(H); the right figure depicts the posterior distribution when we observe tails (T).
The result, depicted in Fig. 6.1, is intuitive. Before we observe a coin flip, we consider
any bias equally likely. But after we observe a coin flip we consider some more likely
than others. In particular, if we observe H, our posterior reflects that we consider
larger values of x more likely, while if we observe T, our posterior reflects the opposite.
Example 6.2. Suppose our waiting time for the arrival of a train y is an exponential
random variable with unknown mean 1/x, which we model as a random variable. In
particular,
py |x (y|x) = xe−xy , y > 0.
Moreover, our prior for x is, itself, an exponential random variable, with mean 1/µ,
i.e.,
px (x) = µe−µx , x > 0,
where µ > 0 is given. Then given a particular waiting time y = y, the posterior
1.25
y1 < y 2 < y 3
<latexit sha1_base64="Y9kuANZ7TAeIr2eoEak6Ck8Dv0E=">AAACZXicZVDRTtswFHUztrFsbGVMe+FhYRXSnpqkoIImHpB44REEBSRSRbeuGywc27JvKFGUb9krfBJfsN+YaTOJlitZPjq+5/qeM9KCW4yip5b3ZuXtu/erH/yPn9Y+f2mvf72wqjCUDagSylyNwDLBJRsgR8GutGGQjwS7HN0ePb9f3jFjuZLnWGo2zCGTfMIpoKPS9kaZxsnWQbJVpr3m3knbnagbzSp4DeIGdEhTJ+l663cyVrTImUQqwNrrONI4rMAgp4LVflJYpoHeQsYqnmtlsPb97RfstYMScmaH1cxTHWw7ZhxMlHFHYjBjlyQ42R9WXOoCmaRzxaQQAarg2Wkw5oZRFKUDQA13qwT0BgxQdHksjKpywBtUSth6kb/XYKzb30/mW1fhwDpxmE2nIfS7uzt74Zk2XGa9fpi7T5QN6+p//ssWpeISWWZAOJdTsKUtc9fj0o6Xs30NLnrduN/tn+52Dveb3FfJJvlJfpGY7JFDckxOyIBQUpI/5IE8tv56a9437/u81Ws1mg2yUN6Pf5L0tpQ=</latexit>
1 <latexit sha1_base64="9LorHwkZWji4XgkPNCji73yVf7E=">AAACaXicZVHRTtswFHUDbJCxrWUPm9hLtAqJvTQJoIL2hLSXPYJYAYlU0a3rBgvHtuwbShTyNbzCB/EN/ASmLRItV7J1dHyPfc/xQAtuMYoeG97S8sqHj6tr/qf1z1++Nlsbp1YVhrIeVUKZ8wFYJrhkPeQo2Lk2DPKBYGeDq78v52fXzFiu5H8sNevnkEk+4hTQUWnzu06rxFzf3LqtrLcT62CZ7v5Om+2oE00qeA/iGWiTWR2lrcafZKhokTOJVIC1F3GksV+BQU4Fq/2ksEwDvYKMVTzXymDt+1tv2AsHJeTM9quJrzrYcswwGCnjlsRgwi5IcHTQr7jUBTJJp4pRIQJUwYvbYMgNoyhKB4Aa7kYJ6CUYoOgymbuqygEvUSlh63n+RoOxbn4/mU5dhT3rxGE2HofQ7ezt7ocn2nCZ7XTD3D2ibFhXr3+waFEqLpFlBoRzOQZb2jJ3PS7teDHb9+B0pxN3O93jvfbhwSz3VfKT/CLbJCb75JD8I0ekRyipyR25Jw+NJ6/l/fA2p61eY6b5RubKaz8DqJu6gw==</latexit>
px|y (x|y3 )
0.75
<latexit sha1_base64="B+2w5UfChOzgI28vdchMdKaTFDg=">AAACaXicZVHRTtswFHUzGCxjo90emOAlWoXEXpqkoIL2hMTLHplGAYlU0a3rBgvHtuwbShTyNbzCB/EN/MRM20lruZKto+N77HuOh1pwi1H03PDeray+X1v/4H/c+PR5s9n6cm5VYSjrUyWUuRyCZYJL1keOgl1qwyAfCnYxvDl5Pb+4ZcZyJc+w1GyQQyb5mFNAR6XNLZ1Wibm9u3dbWe8l1sEy7f5Im+2oE00reAviOWiTeZ2mrcbPZKRokTOJVIC1V3GkcVCBQU4Fq/2ksEwDvYGMVTzXymDt+7v/sVcOSsiZHVRTX3Ww65hRMFbGLYnBlF2S4PhoUHGpC2SSzhTjQgSogle3wYgbRlGUDgA13I0S0GswQNFlsnBVlQNeo1LC1ov8nQZj3fx+Mpu6CvvWicNsMgmh1znYPwz/aMNl1u2FuXtE2bCu/v3BskWpuESWGRDO5QRsacvc9bi04+Vs34LzbifudXq/D9rHR/Pc18kO+U72SEwOyTH5RU5Jn1BSkwfySJ4aL17L++Ztz1q9xlzzlSyU1/4Lpq26gg==</latexit>
0.5
px|y (x|y2 )
<latexit sha1_base64="yK/A8/mxP5h2jLyU5ZsS5msF7YI=">AAACaXicZVHRTtswFHUDDJbBaNnDJvYSrUKClyYBVNCekPayRxArIJEqunXdYOHYln1DibJ8zV7hg/gGfmKmLRItV7J1dHyPfc/xQAtuMYqeGt7S8sqH1bWP/qf1jc+bzdbWhVWFoaxHlVDmagCWCS5ZDzkKdqUNg3wg2OXg9tfL+eUdM5Yr+QdLzfo5ZJKPOAV0VNr8qtMqMXf3f91W1ruJdbBM47202Y460aSC9yCegTaZ1WnaavxMhooWOZNIBVh7HUca+xUY5FSw2k8KyzTQW8hYxXOtDNa+v/OGvXZQQs5sv5r4qoMdxwyDkTJuSQwm7IIER8f9iktdIJN0qhgVIkAVvLgNhtwwiqJ0AKjhbpSA3oABii6TuauqHPAGlRK2nufvNRjr5veT6dRV2LNOHGbjcQjdzuHBUXiuDZfZfjfM3SPKhnX1+geLFqXiEllmQDiXY7ClLXPX49KOF7N9Dy72O3G30z07bJ8cz3JfI9/JD7JLYnJETshvckp6hJKa/CMP5LHx7LW8b972tNVrzDRfyFx57f+kv7qB</latexit>
0
0 1 2 3 4
x
<latexit sha1_base64="sJTJkvFPUj/Ss6xMa2x/8n3ye+g=">AAACUXicZVBNTxsxFHy79IMu/YBy7GXVCKmn7C6ggHpC6qVHUBtAIhF6cV4WC69t2W8botX+gl7LL+upP6U3TJJKTXiSpdH4jT0zI6uk5zz/E8Ubz56/eLn5Ktl6/ebtu+2d9+fe1E5QXxhl3OUIPSmpqc+SFV1aR1iNFF2Mbr883l/8IOel0d95ZmlYYanlRArkQJ3dXW938m4+n/QpKJagA8s5vd6JPg/GRtQVaRYKvb8qcsvDBh1LoahNBrUni+IWS2pkZY3jNkn2/mOvAtRYkR82c/9tuheYcToxLhzN6Zxdk/DkeNhIbWsmLRaKSa1SNuljqnQsHQlWswBQOBmspOIGHQoO2VeeairkGzZG+XaVv7PofPCfDBaum6zvgzgrp9MMe93Dg6Psm3VSl/u9rAqfGJ+1zb+u1yNqIzVT6VCFlFP0Mz+rwk5ou1jv9ik43+8WvW7v7LBzcrzsfRM+wEf4BAUcwQl8hVPogwCCn/AL7qPf0d8Y4nixGkdLzS6sTLz1AJrIsZw=</latexit>
Figure 6.2: The (second-order) Erlang posterior distribution in Example 6.2 for the
exponential distribution parameter x given data y . Progressively larger values of y
concentrate more probability at smaller values of x.
distribution for x is
Example 6.3. Suppose we observe a Gaussian random variable with unknown mean,
35
A second-order Erlang distributed random variable u has PDF
p(u; λ) = λ2 ue−λu , u ∈ U = R+ .
1
y >0
<latexit sha1_base64="vgBQy52/MgGSorKOjCeJMW16hqg=">AAACV3icZVDRTtswFHXCBl3YBmWPvIRVSDw1CaBS7QFV2guPTKwUiVTo1nWDVce27Bu6KMpH8Apfxtdspi3SWo5k6ej4HvueM9KCW4zjF8/f+PBxc6vxKdj+/OXrzm5z79qqwlDWp0ooczMCywSXrI8cBbvRhkE+Emwwmv58vR88MGO5kr+x1GyYQyb5hFNAJw3K9OA8PYjvdltxO54jfE+SJWmRJS7vmt6PdKxokTOJVIC1t0mscViBQU4Fq4O0sEwDnULGKp5rZbAOgsP/1FtHJeTMDqt5ijo8dMo4nCjjjsRwrq5ZcNIdVlzqApmkC8ekECGq8DVbOOaGURSlI0ANd6uE9B4MUHQNrDxV5YD3qJSw9ar+R4Oxbv8gXWxdRX3rzFE2m0XQaZ+enEVX2nCZHXei3H2ibFRXb42vR5SKS2SZAeFSzsCWtszdjGs7We/2Pbk+bieddufXaavXXfbeIPvkOzkiCTkjPXJBLkmfUDIlj+SJPHsv3l9/028sRn1v6flGVuA3/wE43rNB</latexit>
0.8
<latexit sha1_base64="ZeZiJkZkAta4BPKRoRFrLC9gz00=">AAACanicZVHLTuMwFHXDPCDzKrBhxCaaDhKzaRJABbFCYsOSEVNAIlV167rBwrEt+6YlCuFr2ML/8A98BO4DaVquZOvo+B77nuOeFtxiFD3XvKUPHz99Xl7xv3z99v1HfXXt3KrcUNamSihz2QPLBJesjRwFu9SGQdYT7KJ3czw+vxgyY7mS/7DQrJNBKvmAU0BHdesbulsmZnh757ai2k7sGNph8adbb0TNaFLBexDPQIPM6rS7WjtM+ormGZNIBVh7FUcaOyUY5FSwyk9yyzTQG0hZyTOtDFa+v/Ufe+WghIzZTjkxVgVbjukHA2XckhhM2AUJDg46JZc6RybpVDHIRYAqGNsN+twwiqJwAKjhbpSAXoMBii6UuavKDPAalRK2mudvNRjr5veT6dRl2LZOHKajUQit5t7ufnimDZfpTivM3CPKhlX59gmLFqXiEllqQDiXI7CFLTLX49KOF7N9D853mnGr2fq71zg6mOW+TDbJL7JNYrJPjsgJOSVtQsk9eSCP5Kn24q15P73NaatXm2nWyVx5v18BMPe7QA==</latexit>
0.4
<latexit sha1_base64="zTgg4zMfSLorQB9OnAF6MBSSV5s=">AAACXnicZVBNSysxFE3H73l+uxHcDBZBN50ZlSquBDcuFa0KTl+5TdMxmElCcqcfDPM/3Oq/erv3U4xtBVsPBA4n9yT3nLYW3GIU/at4c/MLi0vLK/6f1bX1jc2t7QerckNZgyqhzFMbLBNcsgZyFOxJGwZZW7DH9uvV1/1jjxnLlbzHoWbNDFLJu5wCOumvbhWJ6Q3Kw8T2BketzWpUi0YIfpN4QqpkgpvWVuUi6SiaZ0wiFWDtcxxpbBZgkFPBSj/JLdNAXyFlBc+0Mlj6/sEP9dlRCRmzzWIUpgwOnNIJusq4IzEYqTMW7J43Cy51jkzSsaObiwBV8BUx6HDDKIqhI0ANd6sE9AUMUHRFTD1VZIAvqJSw5bQ+0GCs299PxlsXYcM6c5j2+yHUa6cnZ+GdNlymx/Uwc58oG5bFd/GzEaXiEllqQLiUfbBDO8zcjGs7nu32N3k4rsX1Wv32tHp5Pul9meyRfXJIYnJGLsk1uSENQokhb+SdfFT+e4vemrcxHvUqE88OmYK3+wlTLrY4</latexit>
0.2 px (x)
0
-4 -3 -2 -1 0 1 2 3 4
x<latexit sha1_base64="sJTJkvFPUj/Ss6xMa2x/8n3ye+g=">AAACUXicZVBNTxsxFHy79IMu/YBy7GXVCKmn7C6ggHpC6qVHUBtAIhF6cV4WC69t2W8botX+gl7LL+upP6U3TJJKTXiSpdH4jT0zI6uk5zz/E8Ubz56/eLn5Ktl6/ebtu+2d9+fe1E5QXxhl3OUIPSmpqc+SFV1aR1iNFF2Mbr883l/8IOel0d95ZmlYYanlRArkQJ3dXW938m4+n/QpKJagA8s5vd6JPg/GRtQVaRYKvb8qcsvDBh1LoahNBrUni+IWS2pkZY3jNkn2/mOvAtRYkR82c/9tuheYcToxLhzN6Zxdk/DkeNhIbWsmLRaKSa1SNuljqnQsHQlWswBQOBmspOIGHQoO2VeeairkGzZG+XaVv7PofPCfDBaum6zvgzgrp9MMe93Dg6Psm3VSl/u9rAqfGJ+1zb+u1yNqIzVT6VCFlFP0Mz+rwk5ou1jv9ik43+8WvW7v7LBzcrzsfRM+wEf4BAUcwQl8hVPogwCCn/AL7qPf0d8Y4nixGkdLzS6sTLz1AJrIsZw=</latexit>
Figure 6.3: The prior and posterior distributions for the unknown mean x of a Gaussian
random variable y , using a Gaussian prior as described in Example 6.3. As depicted, for
an observation y > 0, the posterior distribution is a Gaussian distribution with positive
mean, and more concentrated.
where σ 2 and σ02 are given. Then after some straightforward analysis, we obtain36
σ02 σ 2 σ02
px|y (·|y) = N 2 y,
σ +σ02 σ 2 +σ02
In this case, we see that the prior is symmetric—it equally favors positive and negative
values for the unknown mean x. However, given a positive data value y, the revised
belief favors a positive mean, and with increased confidence in the sense that
σ 2 σ02 2
var[x|y = y] = 2 2 ≤ σ0 = var[x],
σ +σ0
36
The required computations are considerably less cumbersome to carry out using valuable forth-
coming insight.
where, as in the case of hypothesis testing, the objective function is referred to as the
Bayes risk.
Analogous to the case of hypothesis testing, the expectation in (6.1) is over x
and y jointly, and hence x̂(·) is a function that minimizes the cost averaged over all
possible (x, y) pairs. Moreover, also analogous to the case of Bayesian hypothesis
testing, solving for the optimum function x̂(·) in (6.1) can, again, be accomplished on
a pointwise basis, i.e., for each particular value y that is observed, we find the best
possible choice (in the sense of (6.1)) for the corresponding estimate x̂(y). Indeed, we
rewrite our objective function in (6.1) in the form
Z +∞ Z +∞
E C(x, f (y)) = C(x, f (y)) px,y (x, y) dx dy
−∞ −∞
Z +∞ Z +∞
= C(x, f (y)) px|y (x|y) dx py (y) dy, (6.2)
−∞ −∞
37
The revised belief (i.e., posterior) is sometimes referred to as a “soft” decision, and an estimate
is correspondingly referred to as a “hard” decision. However, such terminology is not universal.
and note that since py (y) ≥ 0, we minimize (6.2) if we choose x̂(y) to minimize the
term in brackets for each particular value of y, i.e.,
Z +∞
x̂(y) = arg min C(x, a) px|y (x|y) dx. (6.3)
a −∞
for some Ci (·, ·), and with xi and x̂i denoting the ith elements of the vectors x and x̂,
respectively. When (6.4) is satisfied, estimating vector variables x can be handled in a
component-wise manner.
Claim 6.1. The MAE estimate is the median of the posterior belief px|y (·|y), i.e.,
x̂MAE (y) = median px|y (·|y) .
Differentiating the quantity inside braces in (6.6) with respect to a gives, via Leibnitz’
rule, the condition
Z a Z +∞
px|y (x|y) dx − px|y (x|y) dx = 0. (6.7)
−∞ a a=x̂MAE (y)
From (6.8) we see that the x̂MAE (y) is the threshold in x of the posterior belief px|y (x|y)
for which half the probability is located above the threshold and, hence, half is also
below the threshold. Hence, the MAE estimator for x given y = y is the median of
the posterior belief.
A natural generalization of the absolute error cost (6.5) to the case of vector
parameters is
C(a, â) = ∥a − â∥1 , (6.9)
where ∥·∥1 is the L1 norm, defined via
X
∥b∥1 = |bi |,
i
with bi denoting the ith element of b. This is an example of additive cost function,
i.e., satisfies (6.4), and is commonly referred to as the L1 loss.
which uniformly penalizes all estimation errors with magnitude bigger than ϵ.
Claim 6.2. In the limit as ϵ → 0, the MUC estimate is the mode of the posterior
x̂MAP (y) = mode px|y (·|y) = arg max px|y (a|y). (6.11)
a
From this claim, we see that the MAP estimator can be viewed as resulting from a
Bayes cost formulation in which all errors are, in an appropriate sense, equally bad.
Proof. Substituting (6.10) into (6.3) we obtain that the MUC estimator satisfies
Z a+ϵ
ϵ
x̂MUC (y) = arg min 1 − px|y (x|y) dx
a a−ϵ
Z a+ϵ
= arg max px|y (x|y) dx. (6.12)
a a−ϵ
Note that via (6.12) we see that x̂ϵMUC (y) corresponds to the value of a that makes
P(|x − x̂ϵMUC (y)| < ϵ | y = y) as large as possible. This means finding the interval of
length 2ϵ where the posterior density px|y (x|y) is most concentrated.
If we carry this perspective a little further, we see that if we let ϵ get sufficiently
small then the x̂ϵMUC (y) approaches the point corresponding to the peak of the posterior
density, i.e., the MAP estimate (6.11):
lim x̂ϵMUC (y) = arg max px|y (a|y) = x̂MAP (y). (6.13)
ϵ→0 a
The generalization of the limiting minimum uniform cost to the case of vector
parameters corresponds to a cost function that is commonly referred to as the L0
loss.39
from which we obtain what is termed the Bayes least-squares (BLS) estimator. Since
this estimator minimizes the mean-square estimation error, it is often alternatively
referred to as the minimum mean-square error (MMSE) estimator and denoted using
x̂MMSE (·).
Claim 6.3. The BLS (or MMSE) estimate is the mean of the posterior belief px|y (·|y),
i.e.,
x̂BLS (y) = E[x|y = y].
We begin with the simpler case of scalar estimation, for which (6.15) becomes
Z +∞
x̂BLS (y) = arg min (x − a)2 px|y (x|y) dx. (6.16)
a −∞
∂2
Z +∞ Z +∞
2
(x − a) px|y (x|y) dx = 2 px|y (x|y) dx = 2 > 0.
∂a2 −∞ −∞
Hence,
x̂BLS (y) = E[x|y = y], (6.19)
i.e., that the BLS or MMSE estimate of x given y = y is the mean of the posterior
belief px|y (·|y).
When x is a vector, it suffices to note that since the cost criterion (6.14) is additive
[cf.(6.4)], the minimum is achieved by minimizing the mean-square estimation error in
each scalar component. Hence, we obtain
from which we see that the BLS estimate of x given y = y is in general the mean of
the posterior belief px|y (·|y).
The error
e(x, y) = x̂(y) − x, (6.21)
of any estimator x̂(y) can be expressed in the form
e(x, y) = b + e(x, y) − b , (6.22)
where Z +∞ Z +∞
b = E e(x, y) = x̂(y) − x px,y (x, y) dx dy (6.23)
−∞ −∞
is referred to as the (global) bias in the estimator, and where the term in brackets has
zero-mean.
Moreover, the mean-square
T estimation
error in such an estimate is the trace of the
error correlation matrix E ee , which can be expressed via (6.23) in the form
E eeT = Λe + bbT ,
(6.24)
where T
Λe = E e(x, y) − b e(x, y) − b (6.25)
is the error covariance matrix associated with the estimator. Hence we see that in
general, both the bias and covariance contribute to the error correlation and, in turn,
mean-square estimation error.
For the case of a Bayes least-squares estimate specifically, it is straightforward to
verify that there is no bias component to the error, as the following claim asserts.
where the last equality follows from a simple application of the law of iterated
expectation.
That the BLS estimate is unbiased means that the MSE performance of the
estimator is simply the trace of the error covariance matrix ΛBLS . Hence, we have the
following performance characterization.
Claim 6.5. The error covariance matrix is the expected covariance of the posterior
belief, i.e.,42
ΛBLS = E Λx|y (y) . (6.27)
Proof. Using (6.21), (6.25), and (6.26) we obtain that the associated error covariance
is given by T
ΛBLS ≜ Λe = E eeT = E x − E[x|y] x − E[x|y] ,
(6.28)
where we emphasize that the notation ΛBLS is used to refer to the error covariance
of the BLS estimator. Applying iterated expectation to (6.28) we see that the error
covariance can be written as
h T i
ΛBLS = E E x − E[x|y] x − E[x|y] y . (6.29)
However, the inner expectation in (6.29) is simply the covariance of the posterior
belief, i.e., Λx|y , whence (6.27).
E x̂(y) − x g(y)T = 0.
(6.30)
42
Note that given an observed value of y, this posterior covariance Λx|y=y is in general a function
of y. As such, we’ll sometimes use the alternative notation Λx|y (y) for this covariance.
and note, using the law of iterated expectation, that the left-hand side of (6.31) can
in turn be expressed in the form
h i
E x g(y) = E E x g(y) y = E E[x|y] g(y)T .
T T
(6.32)
To prove the “only if” statement, choose x̂(·) = x̂BLS (·) in (6.31), and note that in
this case the right-hand expressions in both (6.32) and (6.31) are identical, verifying
(6.30).
To prove the converse, let us rewrite (6.30) using (6.31) and (6.32) as
Then, since (6.33) must hold for all g(·), let us choose g(y) = E[x|y] − x̂(y) where x̂(·)
is our estimator. In this case (6.33) becomes
T
E E[x|y] − x̂(y) E[x|y] − x̂(y) = 0,
sought, since the latter is much less sensitive to such outliers. One such compromise
is represented by the Huber loss, for which
(
(x − x̂)2 /2 |x − x̂| ≤ δ
Cδ (x, x̂) ≜ (6.34)
δ(|x − x̂| − δ/2) |x − x̂| > δ,
6.7 Examples
Example 6.4 (Revisiting Example 6.1). In this coin toss example, we have, using a
little geometry to calculate the median,
( √
1/ 2 y=H
x̂MAE (y) = √
1 − 1/ 2 y = T.
And clearly (
1 y=H
x̂MAP (y) =
0 y = T.
Finally,
(
Z 1
1y=H 1y=T 2/3 y = H
x̂BLS (y) = E[x|y = y] = x (2x) (2 − 2x) dx =
0 1/3 y = T.
Example 6.5 (Revisiting Example 6.2). In the train waiting time example, we have
1
x̂MAP = , (6.35)
µ+y
since the mode of a second-order Erlang distribution with parameter λ is 1/λ, while
2
x̂BLS (y) =
µ+y
since the mean of a second-order Erlang distribution with parameter λ is 2/λ. Finally,
we have
c
x̂MAE (y) = ,
µ+y
where the constant c is an involved calculation (and well beyond our scope of interest).
Indeed, the median of an Erlang distribution is c/λ, where
1
c = −W−1 − − 1, (6.36)
2e
with W−1 (·) denoting the so-called negative branch of the Lambert W function. Usefully,
it can be shown44 that
5
1.67 ≈ < c ≤ 1 + ln(2) ≈ 1.69,
3
from which we conclude that x̂MAP (y) < x̂MAE (y) < x̂BLS (y).
σ02
x̂MAE (y) = x̂MAP (y) = x̂BLS (y) = y, (6.37)
σ 2 +σ02
due to the fact that the Gaussian distribution is both symmetric and unimodal.
as are developed in Section 7, and neural networks, which often use layering to combine
linear processing with so-called nonlinear activation functions in a multi-layer structure.
In modern deep neural networks, θ can include billions of parameters.
In conjunction with choosing an architecture, one also chooses a cost function
C(·, ·). Popular choices include those described earlier, among others. Still others we
will develop axiomatically in subsequent sections.
ΛBLS ⪯ Λe , (6.38)
Then since a B(p) random variable has variance p(1 − p) it follows that
and in turn
1 1 1 1
λx|y = py (0) λx|y (0) + py (1) λx|y (1) = ·0+ · = .
2 2 4 8
Moreover,
(
3/4 x = 0
px (x) = py (0) px|y (x|0) + py (1) px|y (x|1) =
1/4 x = 1,
whence
3 1 3
λx = · = .
4 4 16
As a result, we have the ordering
From Theorem 6.1 we get that the last two terms in (6.43) are zero. Using this
together with the definition of ΛBLS we get
The right-hand side of (6.44) is in general positive semidefinite, which verifies (6.38),
and equal to zero if and only if g(y) = 0, which using (6.42) yields (6.39).